12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Task automation and instructional planning support with large language models: a systematic review

Giovanni Luna Chontal, Roberto Ángel Meléndez-Armenta, Edgar Degante-Aguilar, Francisco Javier Fernandez-Dominguez · Frontiers in Education · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

7/10
Relevance
3/4
Quality (LMQS)
E
Evidence
2
Citations
39.11
FWCI
Top 10%
Impact

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.3389/feduc.2026.1733861

Methodology & findings

Study design

Systematic review following PRISMA 2020 guidelines.

Sample

N = 16, 2 groups

Primary method

Qualitative synthesis approach. No meta-analysis conducted. Risk of bias assessment using ROBINS-I (ordinal rating scale: low, moderate, serious, or critical risk across seven domains) and CASP checklists (qualitative scoring). Two-reviewer independent assessment with consensus resolution procedures or fourth-reviewer arbitration. Descriptive frequency tallies and narrative synthesis of findings organized by outcome domains (RQ1: efficiency/material quality; RQ2: automation/planning support). Thematic analysis of qualitative findings from included studies.

Main result

The reviewed studies suggest that "LLM use was associated with reported time savings and perceived gains in clarity or usefulness of generated educational resources." However, the authors note that "outcomes and measures were heterogeneous and often self-reported, several risk-of-bias domains were rated as unclear, and evidence was concentrated in higher-education settings with small samples, limiting comparability and causal inference."

Reports effect sizes.

Research paradigm

Critical realist; mixed-methods evidence synthesis with quantitative and qualitative integration

Author conclusions

"Regarding RQ1, the reviewed studies suggest that using LLMs may reduce the time teachers devote to generating educational materials and may improve perceived quality in terms of clarity, coherence, and appropriateness. However, reported time-related benefits varied substantially across contexts and measurement approaches, and many outcomes relied on self-report rather than objective time-on-task measures. As for RQ2, the evidence indicates that LLMs can support automation or task assistance for pedagogical planning activities — lesson planning, schedules, rubrics, and classroom routines, potentially enabling more time for higher-value teacher–student interaction. Nevertheless, effectiveness depends heavily on prompt quality, implementation design, and the degree of technological integration within institutional settings." The authors further conclude: "Recognizing and balancing these factors is fundamental to the sustainable and responsible integration of LLMs in education. These core themes respond directly to the research questions posed, providing a comprehensive perspective on efficiency, perceived quality, and the automation of teachers' work when LLMs are incorporated into educational settings."

Risk of bias

Selection bias: Keyword clusters emphasizing efficiency, time, and automation may have underrepresented studies reporting neutral or negative impacts; Database selection bias: Limited to three databases; excluded education-specialist indexing services (ERIC, PsycINFO, Web of Science education collections); Measurement bias: Heavy reliance on self-reported outcomes rather than objective metrics for time savings and perceived quality; Unclear risk of bias in ROBINS-I assessment: Five of seven domains rated as 'unclear' in quasi-experimental studies (Gasaymeh & AlMohtadi 2024; Winder et al. 2024); Recruitment and reflexivity concerns in qualitative and mixed-methods studies identified in CASP assessment; Publication bias: Studies with positive findings may be more likely to be published; Heterogeneity of implementation factors under-specified across primary studies; Selection bias: Search strategy emphasis on efficiency/time/automation terminology may underrepresent studies with neutral or negative findings; Database coverage bias: Exclusion of ERIC, PsycINFO, and Web of Science education collections may underrepresent K-12 and mainstream education research; Unclear risk of bias in multiple domains: ROBINS-I assessment showed five of seven domains rated as unclear in quasi-experimental studies (Gasaymeh & AlMohtadi 2024; Winder et al. 2024); Recruitment and reflexivity concerns: Noted in qualitative and mixed-methods studies; Self-report bias: Many outcomes rely on self-reported time savings and perceived quality rather than objective measures; Publication bias: No formal assessment of publication bias reported; Temporal confounding: Fast-changing LLM capabilities limit applicability of findings beyond 2023-2025 period; Selection bias: Emphasis on efficiency/automation terminology may underrepresent neutral or negative findings; Database selection bias: Exclusion of ERIC, PsycINFO, and WoS education collections underrepresents K-12 and mainstream education research; Measurement bias: Heavy reliance on self-report and perception-based measures rather than objective metrics; Risk of bias in included studies: Five of seven ROBINS-I domains rated as 'unclear' in quasi-experimental studies (Gasaymeh & AlMohtadi 2024; Winder et al. 2024); Recruitment and reflexivity concerns in qualitative and mixed-methods studies (10 of 13 primary studies); Publication bias: Studies reporting neutral/negative outcomes may be underrepresented

Limitations

  • "Evidence base remains relatively small and heterogeneous, with substantial variation in study design, implementation settings, and outcome definitions, which limits direct comparability across studies." "Most included studies focus on higher education contexts and short-term deployments, constraining generalizability to other levels and to long-term adoption." "Several reported outcomes — time savings and perceived quality — are frequently measured via self-report or perception-based instruments rather than standardized, objective metrics, which increases uncertainty about effect magnitude." "Database coverage was limited to Scopus, ACM Digital Library, and Dimensions
  • consequently, education-specialist indexing services (e.g., ERIC, PsycINFO, and Web of Science education collections) were not searched, which may underrepresent mainstream education journals and pre-tertiary evidence." "Because our research questions and keyword clusters emphasize efficiency, time, and automation, there is a potential selection bias toward studies framed around productivity gains
  • relevant work reporting neutral or negative impacts — or focusing on other pedagogical outcomes without those terms — may be underrepresented."

Open questions raised

  • Standardized and transparent reporting of outcome measures (including objective time-use metrics and validated quality rubrics)
  • Stronger comparative designs (controlled field studies, quasi-experiments, and replication studies across institutions)
  • Broader coverage of educational levels, regions, and resource-constrained settings
  • Implementation research disentangling model-level effects from platform/workflow factors (prompting support, interface design, teacher training)
  • More consistent evaluation of governance issues (privacy, academic integrity, bias mitigation, cost) to support responsible and scalable deployment
  • Longitudinal, classroom-based studies in underrepresented contexts (pre-tertiary, rural, low-resource, multilingual settings)
Data: The original contributions presented in the study are included in the article/supplementary material (Supplementary Table S1 contains complete extraction table). No external datasets with URLs provided.; No datasets available. The authors state: "The original contributions presented in the study are included in the article/Supplementary material, further inquiries can be directed to the corresponding author/s." Supplementary Table S1 with complete extraction data is referenced.; No primary datasets from this review. Supplementary Table S1 (full extraction table) provided within the article/Supplementary material. Authors state: "The original contributions presented in the study are included in the article/Supplementary material, further inquiries can be directed to the corresponding author/s."Code: No code repositories mentioned.; None mentioned.Extracted from: pdfAgreement 49%

Explore related topics

Related papers