12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Open-Source Large Language Models in Education: A Narrative Review of Evidence, Pedagogical Roles, and Learning Outcomes

Michael Pin-Chuan Lin, Jing-Yuan Huang, Daniel Chang, Gerald Tembrevilla, G. Michael Bowen, Eric Poitras et al. · AI in Education · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

7/10
Relevance
1/4
Quality (LMQS)
I
Evidence
0
Citations
0.00
FWCI

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.3390/aieduc2010004

Methodology & findings

Study design

Narrative review of peer-reviewed, human empirical studies examining open-source LLMs in education.

Primary method

Interpretive synthesis of qualitative and mixed-methods empirical studies. No quantitative synthesis or statistical pooling reported.

Main result

The review found that "learner perceptions are generally positive, but evidence linking open-source LLM use to measurable learning outcomes remains emerging and inconsistent." The study synthesized how open-source LLMs are deployed primarily in higher education within computer science and programming domains, with applications focused on post-class tutoring, guidance, and formative feedback.

Reports effect sizes.

Research paradigm

Interpretivism/Hermeneutics

Author conclusions

The authors conclude that "this review maps recurring orchestration dimensions, decision points, and tensions that characterize early implementations, and it proposes a minimal orchestration reporting scaffold (configuration, boundaries, logging, adjudication) intended to support auditability and cross-study comparison as the empirical base develops." They also articulate "a four-role model—Designer, Facilitator, Monitor, and Evaluator—that captures how teacher agency is enacted across AI-supported instructional workflows."

Risk of bias

Publication bias (narrative review without systematic protocol); Selection bias in literature identification (scope limited to peer-reviewed studies only); Domain concentration bias (overrepresentation of computer science and programming); Educational level bias (concentration in higher education); Outcome measurement bias (reliance on reported learner perceptions over objective measures); Publication bias (only peer-reviewed studies included); Selection bias (concentrated in higher education and computer science domains); Geographic/language bias (potential English-language publication bias); Context-specific bias (applications concentrated in post-class tutoring and formative feedback); Selection bias in reviewed studies (concentration in higher education and computer science); Heterogeneity in study designs and outcome measures; Limited diversity in educational contexts and disciplines

Limitations

  • The authors note that "empirical research on how open-source LLMs are deployed in education and what evidence currently supports their integration remains limited and fragmented." Additionally, the review indicates that the "reviewed literature is concentrated in higher education, particularly within computer science and programming domains," which limits generalizability across educational contexts and subject areas.

Open questions raised

  • Limited empirical evidence linking open-source LLM use to measurable learning outcomes
  • Fragmented research base requiring standardization of reporting
  • Need for cross-study comparison frameworks
  • Underrepresentation of research outside higher education and computer science domains
  • Need for better documentation of orchestration dimensions in AI-supported instructional workflows
  • Inconsistent reporting across studies
Data: not_statedCode: not_statedExtracted from: pdfAgreement 61%

Explore related topics

Related papers