12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Does ChatGPT enhance student learning? A systematic review and meta-analysis of experimental studies

Ruiqi Deng, Yuyan Lu, Shasha Liu · Computers & Education · 2024

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

7/10
Relevance
0/4
Quality (LMQS)
E
Evidence
295
Citations
103.58
FWCI
Top 10%
Impact

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1016/j.compedu.2024.105224

Methodology & findings

Study design

Systematic review with meta-analysis of experimental studies.

Sample

Not applicable for meta-analysis of 69 studies, 1 group

Primary method

Meta-analytic methods applied to experimental studies, though specific statistical software and heterogeneity models not detailed in abstract. Authors recommend future use of power analysis.

Main result

The review found that "ChatGPT improves academic performance, affective-motivational states, and higher-order thinking propensities; it reduces mental effort and has no significant effect on self-efficacy." The study also highlights that "ChatGPT interventions are predominantly implemented at the university level, cover various subject areas focusing on language education, are integrated into classroom environments as part of regular educational practices, and primarily involve direct student use of ChatGPT."

Reports effect sizes.

Research paradigm

Positivist/Empiricist

Author conclusions

The authors conclude that "This review provides valuable insights for researchers, instructors, and policymakers evaluating the effectiveness of generative AI integration in educational practice" and propose four key recommendations: distinguishing between output quality and intervention effects through more rigorous assessment methods, evaluating long-term impacts to determine sustainability of benefits, prioritizing objective measures for higher-order thinking, and using power analysis for adequate sample sizing.

Risk of bias

Lack of power analysis in included studies; Concerns regarding post-intervention assessment design; Potential novelty effect bias in affective-motivational outcomes; Reliance on subjective assessments of higher-order thinking; Risk of publication bias toward positive findings; Concerns regarding post-intervention assessments; Risk of novelty effect confounding long-term impacts; Reliance on subjective measures for higher-order thinking; Publication bias potential (not explicitly addressed in abstract); Concerns regarding post-intervention assessment validity; Reliance on subjective rather than objective outcome measures; Potential novelty effects not distinguished from sustained learning gains; Lack of proctored assessments in some studies; Publication bias (primarily positive findings reported)

Limitations

  • The authors identified several methodological limitations: "methodological limitations, such as the lack of power analysis and concerns regarding post-intervention assessments, warrant cautious interpretation of results." Additionally, they note that many studies rely on subjective assessments and lack objective measures for evaluating higher-order thinking, and there is insufficient evidence for long-term sustainability of observed effects beyond potential novelty effects.

Open questions raised

  • Need to distinguish between quality of ChatGPT outputs and positive effects of interventions through more complex, project-based assessments and proctored assessments
  • Evaluate long-term impacts to determine whether positive effects on affective-motivational states are sustained or due to novelty effect
  • Prioritize objective measures to complement subjective assessments of higher-order thinking
  • Use power analysis to determine adequate sample sizes and provide reliable effect size estimates
  • Need for long-term impact studies to determine whether positive effects on affective-motivational states are sustained or due to novelty effects
  • Lack of complex, project-based assessments compared to well-defined problems in post-intervention evaluations
Data: not_statedCode: not_statedExtracted from: pdfAgreement 57%

Explore related topics

Related papers