12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

From Pilots to Practices: A Scoping Review of GenAI-Enabled Personalization in Computer Science Education

Iman Reihanian, Yunfei Hou, Qingquan Sun · AI · 2025

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

7/10
Relevance
1/4
Quality (LMQS)
I
Evidence
4
Citations
6.84
FWCI
Top 10%
Impact

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.3390/ai7010006

Methodology & findings

Study design

Scoping review with purposive sampling of 32 studies from 259 records (2023-2025) in higher-education CS contexts.

Sample

N = 32, 1 group

Primary method

Scoping review methodology (systematic search and thematic mapping). No quantitative statistical analysis reported; qualitative synthesis of design patterns and effectiveness signals.

Main result

The review identified five application domains and found that "designs incorporating explanation-first guidance, solution withholding, graduated hint ladders, and artifact grounding (student code, tests, and rubrics) consistently show more positive learning processes than unconstrained chat interfaces." Additionally, "successful implementations share four patterns: context-aware tutoring anchored in student artifacts, multi-level hint structures requiring reflection, composition with traditional CS infrastructure (autograders and rubrics), and human-in-the-loop quality assurance."

Reports effect sizes.

Research paradigm

Empirical evidence synthesis; systematic mapping of mechanisms and effectiveness signals across heterogeneous studies

Author conclusions

"The evidence supports generative AI as a mechanism for precision scaffolding when embedded in exploration-first, audit-ready workflows that preserve productive struggle while scaling personalized support." The authors recommend an "exploration-first adoption framework emphasizing piloting, instrumentation, learning-preserving defaults, and evidence-based scaling," with attention to "four recurrent risks—academic integrity, privacy, bias and equity, and over-reliance."

Risk of bias

Selection bias from purposive sampling (259 records narrowed to 32 studies); Publication bias favoring positive outcomes in emerging GenAI applications; Recency bias due to narrow publication window (2023-2025); Limited representation of equity and underrepresented populations; Selection bias: purposive sampling from 259 records may not represent all available evidence; Publication bias: focus on recent literature (2023-2025) may exclude earlier foundational work; Scope limitation: restricted to higher-education CS contexts, limiting generalizability; Methodological heterogeneity across included studies prevents quantitative synthesis; Publication bias (only peer-reviewed studies from 2023-2025 included); Selection bias (purposive sampling from 259 records may not be exhaustive); Heterogeneity in study designs and outcome measures across included studies; Potential funding bias in CS education research

Limitations

  • "Critical evidence gaps include longitudinal effects on skill retention, comparative evaluations of guardrail designs, equity impacts at scale, and standardized replication metrics." The review is limited to recent studies (2023-2025) and focuses on higher-education CS contexts, which may not generalize to other educational levels or disciplines.

Open questions raised

  • Longitudinal effects on skill retention
  • Comparative evaluations of guardrail designs
  • Equity impacts at scale
  • Standardized replication metrics
  • Evidence on academic integrity, privacy, bias and equity, and over-reliance risks
  • Operationalization of four recurrent risks: academic integrity, privacy, bias and equity, and over-reliance
Data: not_statedCode: not_statedExtracted from: pdfAgreement 61%

Explore related topics

Related papers