12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Generative AI in mathematics education: Considerations for academic integrity and assessment strategies

Kunti Robiatul Mahmudah, Nur Robiah Nofikusumawati Peni, Faida Musa'ad, Soth Chea, Sommay Shingphachanh · Jurnal Elemen · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
2/4
Quality (LMQS)
I
Evidence
0
Citations
0.00
FWCI

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.29408/jel.v12i2.33851

Methodology & findings

Study design

Systematic review following PRISMA 2020 guidelines.

Sample

N = 18, 1 group

Primary method

Thematic analysis using inductive and deductive coding approaches. Two reviewers independently coded all included studies using open coding. Discrepancies in coding were resolved through consensus. Themes were refined iteratively. Studies were categorized by methodological approach (empirical, conceptual, design-based research) to avoid treating all evidence as methodologically equivalent.

Main result

The review identified that "GenAI poses both risks and opportunities for education. While it challenges longstanding norms around authorship, originality, and effort, it also invites educators and institutions to revisit the purposes and practices of assessment itself." The study found that "traditional forms of assessment, particularly closed-book, product-oriented examinations are increasingly vulnerable to misuse or automation via AI tools" and that "researchers advocate for the design and implementation of authentic assessments that emphasize process, reflection, critical thinking, and originality."

Reports effect sizes.

Research paradigm

Mixed (qualitative and quantitative empirical studies combined with conceptual analysis)

Author conclusions

The authors conclude that "GenAI is reshaping the landscape of education assessment in profound ways. Based on the current body of literature reviewed, this systematic review demonstrates that while the risks to academic integrity are real, particularly regarding authorship, plagiarism, and superficial learning, there is also significant potential for GenAI to enhance feedback, personalization, and engagement when implemented thoughtfully. The key to addressing GenAI's challenges lies not in prohibiting its use, but in transforming assessment design." They emphasize that "institutions must invest in empirical research, cross-sector collaboration, and sustain professional development" and that "Future research should prioritize robust empirical and longitudinal investigations to validate the effectiveness of proposed assessment redesign strategies across diverse disciplinary and institutional contexts."

Risk of bias

Selection bias: Only English-language publications included; publication bias toward positive findings likely in emerging field; Methodological heterogeneity: Mix of empirical, conceptual, and design-based studies not subject to uniform quality assessment; Small sample of included studies (n=18) reduces representativeness; Predominance of non-experimental designs limits causal inference; Geographic and disciplinary variation in included studies may not represent full spectrum of evidence; Reviewer agreement procedures relied on consensus discussion rather than formal inter-rater reliability statistics; Small number of included studies (n=18) limiting generalizability; Predominance of non-experimental research designs (conceptual and exploratory studies); Potential selection bias from limiting search to English-language publications only; Geographic and disciplinary heterogeneity of reviewed studies; Lack of formal risk-of-bias assessment tool applied uniformly across studies; Screening bias: Only English-language publications included; Database selection bias: Limited to five major databases; grey literature not included; Publication bias: Peer-reviewed articles and conference papers only; potentially favors positive findings; Time-period bias: Publication date range 2022-2025 may not capture earlier foundational work; Methodological heterogeneity: Mix of empirical, conceptual, and design-based studies without standardized quality assessment tool; Small sample bias: Only 18 studies included after screening; limited to derive strong conclusions; Disciplinary bias: While stated as general education, few studies from specific domains like mathematics

Limitations

  • The authors state that "Given the relatively small number of included studies (n = 18) and the predominance of non-experimental research designs, the conclusions drawn from this review should be interpreted cautiously." Additionally, they note that "much of the evidence supporting these approaches is derived from conceptual, exploratory, or small-scale empirical studies." The review also identifies that there is "limited understanding of how GenAI influences assessment equity across diverse student populations, especially with regard to linguistic diversity, cultural differences, and socioeconomic disparities."

Open questions raised

  • Limited empirical research on impact of GenAI on student learning outcomes in discipline-specific contexts
  • Insufficient understanding of how GenAI influences assessment equity across diverse student populations, including linguistic diversity and socioeconomic disparities
  • Lack of research on GenAI in non-textual and practice-based disciplines (art, design, laboratory sciences)
  • Need for evaluation of scalability of AI-integrated formative assessments in large or resource-constrained environments
  • Absence of co-creation of ethical frameworks with students in empirical studies
  • Lack of comparative, cross-cultural studies on variations in students' perceptions and usage patterns across educational systems
Extracted from: pdfAgreement 68%

Explore related topics

Related papers