12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Working with Large Language Models: impacts on performance, efficiency and perceived responsibility

Felix Kares, Leon Hannig, Markus Langer · Behaviour and Information Technology · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
3/4
Quality (LMQS)
E
Evidence
0
Citations
0.00
FWCI

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1080/0144929x.2026.2666279

Methodology & findings

Study design

Two field experiments with between-participants (Study 1) and within-subjects repeated-measures designs (Study 2).

Sample

N = 58, 9 groups

Primary method

Independent and dependent t-tests for between- and within-participant comparisons; paired t-tests for Study 2 repeated measures; exploratory within-subject mediation analysis using difference scores (post - pre); nonparametric bootstrapping with 5,000 resamples for indirect effects; Pearson correlations; intraclass correlation coefficient (ICC) for rater reliability; McDonald's omega (ω) for internal consistency reliability

Main result

Across both studies, participants consistently reported reduced information processing, diminished feelings of control, and lower perceived responsibility for outcomes when completing tasks with ChatGPT compared to completing the same tasks without its assistance. "Information processing, perceived control and responsibility were lower when using ChatGPT compared to programming without it. This suggests that participants might have offloaded cognitive labour to the LLM, which is consistent with prior findings." Additionally, "the control group completed significantly more subtasks than the prompt engineering group and there was no significant difference in the time taken," suggesting that prompt engineering advice may impair rather than enhance performance.

Reports effect sizes and confidence intervals.

Research paradigm

Positivist empiricism with quantitative experimentation

Author conclusions

"Across both studies, the consistent reduction in perceived information processing, control, and responsibility suggests a psychological shift driven by cognitive offloading and increased automation. While LLMs can streamline workflows and improve task performance, as seen in the writing task, they may also reduce users' sense of ownership, engagement, and accountability." The authors further conclude: "our results demonstrate that the impact of LLMs is not uniform but strongly depends on contextual factors such as domain expertise, task complexity, and the nature of user training. As organisations increasingly adopt these technologies, it is vital to move beyond a narrow focus on productivity gains and systematically examine how LLMs transform the cognitive, emotional, and motivational foundations of knowledge work."

Risk of bias

Selection bias: Study 1 participants were self-selected volunteers from a single German industrial company department; Study 2 used convenience sampling of graduate students; Attrition/exclusion in Study 1: 7 failed attention check, 3 requested data not be used, 2 excluded for insufficient programming experience; Single-task design in Study 1: Used Lua programming language unfamiliar to most participants, potentially reducing generalizability; Small sample sizes limit statistical power for detecting smaller effects; Study 1 lacked no-AI baseline condition; pre-post differences may reflect time effects or task engagement changes rather than ChatGPT specifically; Different participant populations (professionals vs. students) across studies may confound findings; Counterbalancing of materials in Study 2 mitigates but does not eliminate task difficulty confounds; Selection bias: Study 1 sample of 40 from 76 initial participants (attrition/exclusion of 53%); Study 2 sample limited to graduate students in psychology seminar; Attrition: 7 participants failed attention check in Study 1, 3 requested data not be used, 2 excluded for lack of programming experience; Confounding variables: Study 1 used unfamiliar Lua language which may not generalize; task type and domain differences between studies; Small sample size: Study 1 N=40, Study 2 N=18, likely underpowered for detecting smaller effects; Practice effects: Study 2 within-subjects design may show order or learning effects despite counterbalancing; Measurement timing effects: Study 1 pre-post design without baseline control for time effects or task engagement shifts; Selection bias: Study 1 used voluntary participation from a single company department; Study 2 used convenience sample of graduate students; Attrition: Study 1 excluded 3 participants who requested data not be used, 7 who failed attention checks, and 2 with no programming experience; Confounding: Study 1 lacked no-ChatGPT baseline condition; time effects and task engagement shifts cannot be ruled out; Small sample sizes limiting statistical power in both studies; Task-specific effects: Use of unfamiliar Lua programming language in Study 1 may not generalize; different task types in Study 2 may have different sensitivity to AI effects

Open questions raised

  • Whether negative effects of prompt engineering persist over extended adaptation periods or diminish as users become more fluent in applying prompting strategies
  • How work experience and experience working with AI tools moderate the effectiveness of prompt engineering techniques
  • Whether findings from software development and academic contexts generalize to other domains
  • Mechanisms underlying reduced responsibility in data analysis tasks (different from writing tasks)
  • Long-term effects of sustained AI collaboration and whether reductions in engagement/responsibility accumulate or stabilize
  • Effective interventions to mitigate cognitive and motivational tradeoffs while maintaining performance benefits
Data: All data, materials, and R scripts available at https://osf.io/g9hpc/?view_only=9c0809ab610d40d1b598234f37bdfa7b; Yes. "All data, material, and R scripts can be found under https://osf.io/g9hpc/?view_only=9c0809ab610d40d1b598234f37bdfa7b"Code: https://osf.io/g9hpc/?view_only=9c0809ab610d40d1b598234f37bdfa7b (Open Science Framework repository with R scripts); https://osf.io/g9hpc/?view_only=9c0809ab610d40d1b598234f37bdfa7b (Open Science Framework); Open Science Framework (OSF): https://osf.io/g9hpc/?view_only=9c0809ab610d40d1b598234f37bdfa7bExtracted from: pdfAgreement 46%

Explore related topics

Related papers