Generative AI in mathematics education: Considerations for academic integrity and assessment strategies
Kunti Robiatul Mahmudah, Nur Robiah Nofikusumawati Peni, Faida Musa'ad, Soth Chea, Sommay Shingphachanh · Jurnal Elemen · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.29408/jel.v12i2.33851
Methodology & findings
Study design
Systematic review following PRISMA 2020 guidelines.
Sample
N = 18, 1 group
Primary method
Thematic analysis using inductive and deductive coding approaches. Two reviewers independently coded all included studies using open coding. Discrepancies in coding were resolved through consensus. Themes were refined iteratively. Studies were categorized by methodological approach (empirical, conceptual, design-based research) to avoid treating all evidence as methodologically equivalent.
Main result
The review identified that "GenAI poses both risks and opportunities for education. While it challenges longstanding norms around authorship, originality, and effort, it also invites educators and institutions to revisit the purposes and practices of assessment itself." The study found that "traditional forms of assessment, particularly closed-book, product-oriented examinations are increasingly vulnerable to misuse or automation via AI tools" and that "researchers advocate for the design and implementation of authentic assessments that emphasize process, reflection, critical thinking, and originality."
Reports effect sizes.
Research paradigm
Mixed (qualitative and quantitative empirical studies combined with conceptual analysis)
Author conclusions
The authors conclude that "GenAI is reshaping the landscape of education assessment in profound ways. Based on the current body of literature reviewed, this systematic review demonstrates that while the risks to academic integrity are real, particularly regarding authorship, plagiarism, and superficial learning, there is also significant potential for GenAI to enhance feedback, personalization, and engagement when implemented thoughtfully. The key to addressing GenAI's challenges lies not in prohibiting its use, but in transforming assessment design." They emphasize that "institutions must invest in empirical research, cross-sector collaboration, and sustain professional development" and that "Future research should prioritize robust empirical and longitudinal investigations to validate the effectiveness of proposed assessment redesign strategies across diverse disciplinary and institutional contexts."
Risk of bias
Selection bias: Only English-language publications included; publication bias toward positive findings likely in emerging field; Methodological heterogeneity: Mix of empirical, conceptual, and design-based studies not subject to uniform quality assessment; Small sample of included studies (n=18) reduces representativeness; Predominance of non-experimental designs limits causal inference; Geographic and disciplinary variation in included studies may not represent full spectrum of evidence; Reviewer agreement procedures relied on consensus discussion rather than formal inter-rater reliability statistics; Small number of included studies (n=18) limiting generalizability; Predominance of non-experimental research designs (conceptual and exploratory studies); Potential selection bias from limiting search to English-language publications only; Geographic and disciplinary heterogeneity of reviewed studies; Lack of formal risk-of-bias assessment tool applied uniformly across studies; Screening bias: Only English-language publications included; Database selection bias: Limited to five major databases; grey literature not included; Publication bias: Peer-reviewed articles and conference papers only; potentially favors positive findings; Time-period bias: Publication date range 2022-2025 may not capture earlier foundational work; Methodological heterogeneity: Mix of empirical, conceptual, and design-based studies without standardized quality assessment tool; Small sample bias: Only 18 studies included after screening; limited to derive strong conclusions; Disciplinary bias: While stated as general education, few studies from specific domains like mathematics
Limitations
- The authors state that "Given the relatively small number of included studies (n = 18) and the predominance of non-experimental research designs, the conclusions drawn from this review should be interpreted cautiously." Additionally, they note that "much of the evidence supporting these approaches is derived from conceptual, exploratory, or small-scale empirical studies." The review also identifies that there is "limited understanding of how GenAI influences assessment equity across diverse student populations, especially with regard to linguistic diversity, cultural differences, and socioeconomic disparities."
Open questions raised
- Limited empirical research on impact of GenAI on student learning outcomes in discipline-specific contexts
- Insufficient understanding of how GenAI influences assessment equity across diverse student populations, including linguistic diversity and socioeconomic disparities
- Lack of research on GenAI in non-textual and practice-based disciplines (art, design, laboratory sciences)
- Need for evaluation of scalability of AI-integrated formative assessments in large or resource-constrained environments
- Absence of co-creation of ethical frameworks with students in empirical studies
- Lack of comparative, cross-cultural studies on variations in students' perceptions and usage patterns across educational systems
Explore related topics
Related papers
- Preferred reporting items for systematic reviews and meta-analyses: the PRISMA statementDavid Moher · 2009 · 83,271 citations
- PRISMA Extension for Scoping Reviews (PRISMA-ScR): Checklist and ExplanationAndrea C. Tricco · 2018 · 40,391 citations
- Cochrane Handbook for Systematic Reviews of Interventions2019 · 14,420 citations
- PRISMA 2020 explanation and elaboration: updated guidance and exemplars for reporting systematic reviewsMatthew J. Page · 2021 · 10,956 citations
- Estimating the reproducibility of psychological scienceAlexander A. Aarts · 2015 · 8,669 citations
- Updated methodological guidance for the conduct of scoping reviewsMicah D.J. Peters · 2020 · 6,688 citations