Working with Large Language Models: impacts on performance, efficiency and perceived responsibility
Felix Kares, Leon Hannig, Markus Langer · Behaviour and Information Technology · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1080/0144929x.2026.2666279
Methodology & findings
Study design
Two field experiments with between-participants (Study 1) and within-subjects repeated-measures designs (Study 2).
Sample
N = 58, 9 groups
Primary method
Independent and dependent t-tests for between- and within-participant comparisons; paired t-tests for Study 2 repeated measures; exploratory within-subject mediation analysis using difference scores (post - pre); nonparametric bootstrapping with 5,000 resamples for indirect effects; Pearson correlations; intraclass correlation coefficient (ICC) for rater reliability; McDonald's omega (ω) for internal consistency reliability
Main result
Across both studies, participants consistently reported reduced information processing, diminished feelings of control, and lower perceived responsibility for outcomes when completing tasks with ChatGPT compared to completing the same tasks without its assistance. "Information processing, perceived control and responsibility were lower when using ChatGPT compared to programming without it. This suggests that participants might have offloaded cognitive labour to the LLM, which is consistent with prior findings." Additionally, "the control group completed significantly more subtasks than the prompt engineering group and there was no significant difference in the time taken," suggesting that prompt engineering advice may impair rather than enhance performance.
Reports effect sizes and confidence intervals.
Research paradigm
Positivist empiricism with quantitative experimentation
Author conclusions
"Across both studies, the consistent reduction in perceived information processing, control, and responsibility suggests a psychological shift driven by cognitive offloading and increased automation. While LLMs can streamline workflows and improve task performance, as seen in the writing task, they may also reduce users' sense of ownership, engagement, and accountability." The authors further conclude: "our results demonstrate that the impact of LLMs is not uniform but strongly depends on contextual factors such as domain expertise, task complexity, and the nature of user training. As organisations increasingly adopt these technologies, it is vital to move beyond a narrow focus on productivity gains and systematically examine how LLMs transform the cognitive, emotional, and motivational foundations of knowledge work."
Risk of bias
Selection bias: Study 1 participants were self-selected volunteers from a single German industrial company department; Study 2 used convenience sampling of graduate students; Attrition/exclusion in Study 1: 7 failed attention check, 3 requested data not be used, 2 excluded for insufficient programming experience; Single-task design in Study 1: Used Lua programming language unfamiliar to most participants, potentially reducing generalizability; Small sample sizes limit statistical power for detecting smaller effects; Study 1 lacked no-AI baseline condition; pre-post differences may reflect time effects or task engagement changes rather than ChatGPT specifically; Different participant populations (professionals vs. students) across studies may confound findings; Counterbalancing of materials in Study 2 mitigates but does not eliminate task difficulty confounds; Selection bias: Study 1 sample of 40 from 76 initial participants (attrition/exclusion of 53%); Study 2 sample limited to graduate students in psychology seminar; Attrition: 7 participants failed attention check in Study 1, 3 requested data not be used, 2 excluded for lack of programming experience; Confounding variables: Study 1 used unfamiliar Lua language which may not generalize; task type and domain differences between studies; Small sample size: Study 1 N=40, Study 2 N=18, likely underpowered for detecting smaller effects; Practice effects: Study 2 within-subjects design may show order or learning effects despite counterbalancing; Measurement timing effects: Study 1 pre-post design without baseline control for time effects or task engagement shifts; Selection bias: Study 1 used voluntary participation from a single company department; Study 2 used convenience sample of graduate students; Attrition: Study 1 excluded 3 participants who requested data not be used, 7 who failed attention checks, and 2 with no programming experience; Confounding: Study 1 lacked no-ChatGPT baseline condition; time effects and task engagement shifts cannot be ruled out; Small sample sizes limiting statistical power in both studies; Task-specific effects: Use of unfamiliar Lua programming language in Study 1 may not generalize; different task types in Study 2 may have different sensitivity to AI effects
Open questions raised
- Whether negative effects of prompt engineering persist over extended adaptation periods or diminish as users become more fluent in applying prompting strategies
- How work experience and experience working with AI tools moderate the effectiveness of prompt engineering techniques
- Whether findings from software development and academic contexts generalize to other domains
- Mechanisms underlying reduced responsibility in data analysis tasks (different from writing tasks)
- Long-term effects of sustained AI collaboration and whether reductions in engagement/responsibility accumulate or stabilize
- Effective interventions to mitigate cognitive and motivational tradeoffs while maintaining performance benefits
Explore related topics
Related papers
- Estimating the reproducibility of psychological scienceAlexander A. Aarts · 2015 · 8,669 citations
- The effects of over-reliance on AI dialogue systems on students' cognitive abilities: a systematic reviewChunpeng Zhai · 2024 · 1,009 citations
- Autonomous chemical research with large language modelsDaniil A. Boiko · 2023 · 809 citations
- To use or not to use ChatGPT in higher education? A study of students’ acceptance and use of technologyArtur Strzelecki · 2023 · 691 citations
- Teacher support and student motivation to learn with Artificial Intelligence (AI) based chatbotThomas K. F. Chiu · 2023 · 617 citations
- Do AI chatbots improve students learning outcomes? Evidence from a meta‐analysisRong Wu · 2023 · 469 citations