12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Reimagining writing assessment for the AI era: a systematic review on balancing AI support and authentic skill growth

Mohamed Sayed Abdellatif, Mohammed A. Alshehri, Ali Lamouchi, Mohammed Rahmath, Mohamed Ali Nemt-allah · Frontiers in Psychology · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
0/4
Quality (LMQS)
E
Evidence
0
Citations
0.00
FWCI

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.3389/fpsyg.2026.1809174

Methodology & findings

Study design

Systematic review conducted in accordance with PRISMA 2020 guidelines.

Sample

Review of 19 studies (not a primary empirical study)

Primary method

Inter-rater reliability assessed using Cohen's kappa coefficient (κ = 0.87). Quality appraisal conducted using design-appropriate critical appraisal tools: Newcastle-Ottawa Scale (quantitative studies), CASP checklist (qualitative studies), MMAT criteria (mixed-methods studies). For individual included studies: descriptive statistics extracted where available; qualitative studies analyzed for themes, conceptual frameworks, and representative quotations. Narrative synthesis approach employed rather than quantitative meta-analysis due to heterogeneity across study designs, settings, populations, and outcome measures. Thematic analysis using deductive coding (organized by five guiding research questions) followed by inductive coding (emergent patterns within themes). Constant comparison applied attending to convergence/divergence by geography, institutional context, stakeholder group, disciplinary setting, and student characteristics.

Main result

The study found that "students deploy AI tools in a stratified, cognitively strategic manner—using generative models such as ChatGPT for higher-order planning and drafting while reserving automated writing evaluation tools for mechanical revision." Additionally, "ChatGPT emerged as the dominant AI platform, utilized by 88.8% of students in certain samples, followed by Grammarly at 67.4%." The review also identifies that "stakeholder perspectives diverge fundamentally: students prioritize accessibility and productivity, faculty foreground integrity and assessment validity, and administrators adopt cautious openness while disproportionately directing guidance toward faculty (67%) over students (17.8%)."

Reports effect sizes.

Research paradigm

Mixed-methods interpretivism with evidence synthesis orientation

Author conclusions

The authors conclude that "the central challenge is not technological but pedagogical and ethical: institutions must move from reactive policing of AI use toward proactive curriculum redesign that positions critical judgment, ethical reasoning, and authentic authorship as the irreducible core of academic writing education." They further state that "Success requires moving beyond binary framings of AI as either threat or panacea toward nuanced integration grounded in equity, focused on authentic learning, and committed to developing the distinctly human capacities for critical judgment, creative synthesis, and ethical reasoning that remain essential regardless of technological advancement." The review emphasizes that "promising innovations are emerging that suggest viable pathways forward: flexible policy architectures that recognize contextual variation, process-oriented assessment that privileges learning over products, and explicit AI literacy curricula that position technological competence as legitimate educational outcome."

Risk of bias

Publication bias: positive-utility reporting bias favoring successful implementations; null/negative findings scarce; Social desirability bias: self-reported data on sensitive misconduct behaviors may distort reporting; Selection bias: large-scale surveys employed rigorous controls, but substantial evidence relies on convenience sampling; Time-lag bias: early success stories receive publication priority while longer-term outcomes remain undocumented; Geographic bias: research concentrated in Global North (North America, Europe, China, East Asia); significant blind spots for Global South contexts; Epistemological asymmetry: frameworks and best practices reflect particular cultural assumptions about authorship that may not translate across diverse educational traditions; Institutional stratification bias: elite universities demonstrate sophisticated integration while under-resourced institutions struggle; Positive-utility reporting bias favoring successful implementations; Convenience sampling in included studies; Self-reported data bias, particularly regarding academic misconduct (social desirability effects); Time-lag bias where early success stories receive publication priority; Geographic concentration bias toward Global North contexts; Systemic marginalization of student voices in favor of institutional integrity narratives; Absence of longitudinal designs limiting causal claims; Publication bias: positive-utility reporting bias favoring successful implementations while null or negative findings remain scarce; Social desirability bias: self-reported data on sensitive topics like academic misconduct may distort reporting; Convenience sampling: studies predominantly employed convenience sampling rather than random sampling; Geographic bias: research concentrated in Global North (North America, Europe, China, East Asia) with limited Global South representation; Selection bias: marginalization of student voices in favor of institutional integrity narratives; Temporal instability: rapid evolution of AI capabilities may render current findings obsolete

Limitations

  • The review's conclusions are "further constrained by the relatively small corpus of 19 included studies, which limits the breadth of evidence and the weight that can be placed on any emerging pattern." The authors note that "the predominance of convenience sampling and reliance on self-reported data introduce potential biases, particularly around sensitive topics like academic misconduct where social desirability effects may distort reporting." Additionally, "the absence of longitudinal research designs means that assertions about long-term impacts on critical thinking, independent problem-solving, and professional competence remain largely speculative rather than empirically grounded." Geographic limitations are noted: "Geographic concentration of research effort creates significant blind spots regarding how AI integration unfolds in diverse institutional and cultural contexts." Finally, "the rapid evolution of AI capabilities introduces temporal instability where findings about current system limitations may quickly become obsolete."

Open questions raised

  • Absence of longitudinal research designs to document long-term impacts on critical thinking, independent problem-solving, and professional competence
  • Limited evidence from Global South institutional and cultural contexts
  • Sparse empirical documentation of how AI integration unfolds in diverse disciplinary settings beyond humanities and STEM
  • Insufficient longitudinal data on critical thinking skills erosion
  • Need for expanded evidence base (currently only 19 studies) to strengthen transferability of findings
  • Temporal instability: rapid evolution of AI capabilities means current findings about system limitations may become obsolete
Extracted from: pdfAgreement 64%

Explore related topics

Related papers