12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Do AI chatbots improve students learning outcomes? Evidence from a meta‐analysis

Rong Wu, Zhonggen Yu · British Journal of Educational Technology · 2023

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

5/10
Relevance
2/4
Quality (LMQS)
E
Evidence
469
Citations
120.62
FWCI
Top 10%
Impact

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1111/bjet.13334

Methodology & findings

Study design

Meta-analysis of 24 randomized controlled and quasi-experimental studies.

Primary method

Meta-analysis utilizing Stata software (version 14). Moderator analyses were conducted for educational levels and intervention duration.

Main result

The results indicated that "AI chatbots had a large effect on students' learning outcomes." Specifically, "using AI chatbots could significantly improve learning outcomes in terms of learning performance (ES = 1.028, 95% CI = [0.580, 1.476]), motivation (ES = 1.020, 95% CI = [0.278, 1.763]), self-efficacy (ES = 1.206, 95% CI = [0.357, 2.055]), interest (ES = 1.084, 95% CI = [0.220, 1.947]), and perceived value of learning (ES = 1.397, 95% CI = [0.228, 2.566])." Additionally, "short interventions with a duration of shorter than ten weeks have a large and significant effect size (ES = 1.179, 95% CI = [0.752, 1.606]), whereas long interventions of a duration of ten weeks or longer yield a small and significant effect (ES = 0.492, 95% CI = [0.046, 0.937])."

Reports effect sizes.

Research paradigm

Positivist/empiricist - quantitative synthesis of experimental evidence

Author conclusions

"The results indicated that AI chatbots had a large effect on students' learning outcomes. Moreover, AI chatbots had a greater effect on students in higher education, compared to those in primary education and secondary education. In addition, short interventions were found to have a stronger effect on students' learning outcomes than long interventions. It could be explained by the argument that the novelty effects of AI chatbots could improve learning outcomes in short interventions, but it has worn off in the long interventions. Future designers and educators should make attempt to increase students' learning outcomes by equipping AI chatbots with human-like avatars, gamification elements and emotional intelligence."

Risk of bias

Publication bias: Egger's test indicated potential publication bias in learning performance, motivation, and interest outcomes; Selection bias: Studies limited to English-language publications; Heterogeneity: Very high I² values (93.9% overall) indicating substantial between-study heterogeneity; Study quality variation: Included studies had varying quality assessment scores; Novelty effect: Short interventions may benefit from novelty effects that wear off in long-term use; Publication bias detected in learning performance (Egger's test p = 0.005); trim-and-fill analysis adjusted effect size to ES = 0.944 with 21 missing studies; High heterogeneity across studies (I² = 94%), suggesting unmeasured moderators or methodological differences; Limited investigation of negative effects of AI chatbots (as acknowledged by authors); Inclusion limited to English-language publications, introducing potential language bias; Quality assessment variability in primary studies (inter-rater reliability 0.93, while acceptable, indicates some disagreement); Not explicitly stated in the abstract. Potential risk factors for meta-analyses could include: publication bias (not mentioned as assessed), heterogeneity among included studies, selective reporting in primary studies, and variation in study quality among the 24 RCTs.; Publication bias (not assessed in abstract); Study heterogeneity (conflicting evidence in source studies); Potential selection bias in included studies

Limitations

  • The authors acknowledge that "there has been still a need to conduct a more comprehensive review" and that "relatively little research has reported and analysed the negative effects of AI chatbots on students' learning outcomes." They note that "there is a notable paucity of empirical research focusing specifically on the use of ChatGPT in education." Additionally, the high heterogeneity in effect sizes (I² = 93.9% overall) suggests substantial variability across studies, and the authors note that "students' learning outcomes might be influenced by the context of AI chatbot use."

Open questions raised

  • Limited research on negative effects of AI chatbots on student learning
  • Paucity of empirical research on ChatGPT specifically in educational contexts
  • Need to examine mechanisms underlying the effects of AI chatbots
  • Gender and cultural differences in AI chatbot-supported learning outcomes not yet explored
  • Need to identify which types of learners benefit most from AI chatbot-supported learning
  • Comparison of effectiveness of different types of AI chatbots needed
Data: "We make sure that all data and materials support our published claims and comply with field standards." However, no specific dataset URLs or repositories are provided.; "We make sure that all data and materials support our published claims and comply with field standards." However, no specific dataset URLs or repositories are provided in the paper.Code: No code repositories mentioned in the paperExtracted from: pdfAgreement 53%

Explore related topics

Related papers