Generative AI in teacher education: a systematic review
Libin Yuan, Rafiza Abdul Razak, Amirrudin Kamsin, Siti-Soraya Abdul-Rahman · International Journal of Evaluation and Research in Education (IJERE) · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.11591/ijere.v15i2.37225
Methodology & findings
Study design
Systematic review following PRISMA guidelines with the following steps: (1) establishing clear inclusion and exclusion criteria; (2) formulating and executing a comprehensive search strategy in WoS and Scopus databases; (3) screening and selecting eligible studies; (4) systematically describing and evaluating the included studies; (5) synthesizing and analyzing the results.
Sample
N = 35, 2 groups
Primary method
Thematic synthesis and narrative analysis of qualitative, quantitative, and mixed-method studies. Two researchers independently screened studies with inter-rater reliability measured using Cohen's kappa (κ=0.85). Publication year analysis and geographical distribution analysis descriptively presented. No inferential statistical tests reported for the systematic review itself.
Main result
The analysis reveals that GenAI has shown "diverse applications in teacher education" including "understanding the perceptions and needs of key stakeholders, supporting instructional resources and content generation, informing curriculum design and development, fostering student-AI collaborative learning, guiding educational practice and policy development, and providing tools for research and evaluation." The field is dominated by "qualitative studies account for the majority (60%), most of them focus on stakeholder perceptions and needs analysis" with "quantitative studies (29%) tend to concentrate on the development of competency assessment tools." Geographically, "China (26%) tends to focus on curriculum design and policy integration" while "the United States (14%) are more oriented toward human-AI collaborative learning and ethical reflection," and "regions such as Africa and Latin America are notably absent."
Reports effect sizes and confidence intervals.
Research paradigm
Mixed methods (predominantly qualitative interpretivism with quantitative descriptive components)
Author conclusions
The authors conclude: "This study examines GenAI's applications in teacher education, analyzing its benefits, challenges, and research gaps, while identifying three critical tensions between research and practice. First, the gap between technological capabilities and educational implementation. While GenAI supports tasks such as lesson planning, many teachers still lack the skills needed to evaluate the pedagogical quality of AI-generated content. Second, the gap between research ecology and global diversity. Current research remains fragmented, with studies disproportionately concentrated in China and the United States. This limits the cross-cultural applicability of GenAI integration and leaves the needs of teachers in resource-constrained regions insufficiently represented. Third, the gap between research stage and practical demands." They further state: "Current conclusions about GenAI's effectiveness are tentative, constrained by methodological limitations and rapid technological evolution."
Risk of bias
Selection bias: Exclusion of non-English publications may overlook innovative practices from non-Anglophone contexts; Sample bias: Limited sample diversity in included studies (single universities, narrow disciplinary focus); Reporting bias: Self-reported perceptions without triangulation through classroom observations; Geographic bias: Severe geographical imbalance with most studies concentrated in China (26%) and United States (14%), with Africa and Latin America notably absent; Temporal bias: Many studies rely on short-term or cross-sectional data; most studies based on outdated AI models; Quality assessment bias: No formal quality assessment was conducted to evaluate potential bias in included studies; Language bias: English-only publications selected; Selection bias: Limitation to English-language publications only may exclude non-Anglophone contexts; Selection bias: Exclusion of K-12 settings and non-peer-reviewed studies; Reporting bias: No formal quality assessment of included studies conducted; Self-reporting bias: Studies rely heavily on self-reported perceptions without triangulation; Publication bias: Recent submission deadline (July 2025) may have resulted in underreporting of 2025 studies; Sample bias: Limited sample representativeness across included studies with narrow participant contexts; Geographical bias: Severe geographical imbalance with overrepresentation of China (26%) and United States (14%), with notable absence of Africa and Latin America; Language bias: exclusion of non-English publications; Selection bias: limited to WoS and Scopus databases only; Geographical bias: severe underrepresentation of Africa and Latin America (26% China, 14% USA); Methodological quality bias: no formal quality assessment of included studies; Publication bias: not formally assessed; Self-reporting bias noted in included studies; Sample representativeness issues in included studies
Limitations
- The authors state: "The limitations of the study include: i) the literature search was limited to publications up to July 2025, potentially omitting recent developments
- ii) the exclusion of non-English publications may have overlooked innovative practices from non-Anglophone contexts
- and iii) no formal quality assessment was conducted to evaluate potential bias in the included studies." Additionally, reviewed studies suffer from "limited sample representativeness, with participants drawn from narrow contexts" such as "Barrett and Pack [21] focus on a single university and mostly English teachers
- [26] survey eight Hong Kong universities concentrated in engineering and science
- and Chan and Tsi [27] examine only two secondary English teachers." Many studies "rely heavily on self-reported perceptions without triangulation through classroom observations or behavioral data, creating a gap between reported attitudes and actual practices."
Open questions raised
- Limited sample diversity and generalizability across regions, school types, and academic disciplines
- Lack of large-scale empirical evidence; most findings remain conceptual without experimental or survey-based validation
- Long-term effects underexplored due to reliance on short-term or cross-sectional data
- Limited exploration of GenAI across science disciplines in pre-service teacher education
- Outdated tools or focus narrowly on text-based AI, overlooking emerging multimodal developments
- Ethical guidelines frequently mentioned but remain vague and lack actionable standards
Explore related topics
Related papers
- What if the devil is my guardian angel: ChatGPT as a case study of using chatbots in educationAhmed Tlili · 2023 · 1,587 citations
- A SWOT analysis of ChatGPT: Implications for educational practice and researchMohammadreza Farrokhnia · 2023 · 1,171 citations
- Shaping the Future of Education: Exploring the Potential and Consequences of AI and ChatGPT in Educational SettingsSimone Grassini · 2023 · 921 citations
- Revolutionizing education with AI: Exploring the transformative potential of ChatGPTTufan Adıgüzel · 2023 · 858 citations
- Autonomous chemical research with large language modelsDaniil A. Boiko · 2023 · 809 citations
- Embracing the future of Artificial Intelligence in the classroom: the relevance of AI literacy, prompt engineering, and critical thinking in modern educationYoshija Walter · 2024 · 805 citations