Task automation and instructional planning support with large language models: a systematic review
Giovanni Luna Chontal, Roberto Ángel Meléndez-Armenta, Edgar Degante-Aguilar, Francisco Javier Fernandez-Dominguez · Frontiers in Education · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.3389/feduc.2026.1733861
Methodology & findings
Study design
Systematic review following PRISMA 2020 guidelines.
Sample
N = 16, 2 groups
Primary method
Qualitative synthesis approach. No meta-analysis conducted. Risk of bias assessment using ROBINS-I (ordinal rating scale: low, moderate, serious, or critical risk across seven domains) and CASP checklists (qualitative scoring). Two-reviewer independent assessment with consensus resolution procedures or fourth-reviewer arbitration. Descriptive frequency tallies and narrative synthesis of findings organized by outcome domains (RQ1: efficiency/material quality; RQ2: automation/planning support). Thematic analysis of qualitative findings from included studies.
Main result
The reviewed studies suggest that "LLM use was associated with reported time savings and perceived gains in clarity or usefulness of generated educational resources." However, the authors note that "outcomes and measures were heterogeneous and often self-reported, several risk-of-bias domains were rated as unclear, and evidence was concentrated in higher-education settings with small samples, limiting comparability and causal inference."
Reports effect sizes.
Research paradigm
Critical realist; mixed-methods evidence synthesis with quantitative and qualitative integration
Author conclusions
"Regarding RQ1, the reviewed studies suggest that using LLMs may reduce the time teachers devote to generating educational materials and may improve perceived quality in terms of clarity, coherence, and appropriateness. However, reported time-related benefits varied substantially across contexts and measurement approaches, and many outcomes relied on self-report rather than objective time-on-task measures. As for RQ2, the evidence indicates that LLMs can support automation or task assistance for pedagogical planning activities — lesson planning, schedules, rubrics, and classroom routines, potentially enabling more time for higher-value teacher–student interaction. Nevertheless, effectiveness depends heavily on prompt quality, implementation design, and the degree of technological integration within institutional settings." The authors further conclude: "Recognizing and balancing these factors is fundamental to the sustainable and responsible integration of LLMs in education. These core themes respond directly to the research questions posed, providing a comprehensive perspective on efficiency, perceived quality, and the automation of teachers' work when LLMs are incorporated into educational settings."
Risk of bias
Selection bias: Keyword clusters emphasizing efficiency, time, and automation may have underrepresented studies reporting neutral or negative impacts; Database selection bias: Limited to three databases; excluded education-specialist indexing services (ERIC, PsycINFO, Web of Science education collections); Measurement bias: Heavy reliance on self-reported outcomes rather than objective metrics for time savings and perceived quality; Unclear risk of bias in ROBINS-I assessment: Five of seven domains rated as 'unclear' in quasi-experimental studies (Gasaymeh & AlMohtadi 2024; Winder et al. 2024); Recruitment and reflexivity concerns in qualitative and mixed-methods studies identified in CASP assessment; Publication bias: Studies with positive findings may be more likely to be published; Heterogeneity of implementation factors under-specified across primary studies; Selection bias: Search strategy emphasis on efficiency/time/automation terminology may underrepresent studies with neutral or negative findings; Database coverage bias: Exclusion of ERIC, PsycINFO, and Web of Science education collections may underrepresent K-12 and mainstream education research; Unclear risk of bias in multiple domains: ROBINS-I assessment showed five of seven domains rated as unclear in quasi-experimental studies (Gasaymeh & AlMohtadi 2024; Winder et al. 2024); Recruitment and reflexivity concerns: Noted in qualitative and mixed-methods studies; Self-report bias: Many outcomes rely on self-reported time savings and perceived quality rather than objective measures; Publication bias: No formal assessment of publication bias reported; Temporal confounding: Fast-changing LLM capabilities limit applicability of findings beyond 2023-2025 period; Selection bias: Emphasis on efficiency/automation terminology may underrepresent neutral or negative findings; Database selection bias: Exclusion of ERIC, PsycINFO, and WoS education collections underrepresents K-12 and mainstream education research; Measurement bias: Heavy reliance on self-report and perception-based measures rather than objective metrics; Risk of bias in included studies: Five of seven ROBINS-I domains rated as 'unclear' in quasi-experimental studies (Gasaymeh & AlMohtadi 2024; Winder et al. 2024); Recruitment and reflexivity concerns in qualitative and mixed-methods studies (10 of 13 primary studies); Publication bias: Studies reporting neutral/negative outcomes may be underrepresented
Limitations
- "Evidence base remains relatively small and heterogeneous, with substantial variation in study design, implementation settings, and outcome definitions, which limits direct comparability across studies." "Most included studies focus on higher education contexts and short-term deployments, constraining generalizability to other levels and to long-term adoption." "Several reported outcomes — time savings and perceived quality — are frequently measured via self-report or perception-based instruments rather than standardized, objective metrics, which increases uncertainty about effect magnitude." "Database coverage was limited to Scopus, ACM Digital Library, and Dimensions
- consequently, education-specialist indexing services (e.g., ERIC, PsycINFO, and Web of Science education collections) were not searched, which may underrepresent mainstream education journals and pre-tertiary evidence." "Because our research questions and keyword clusters emphasize efficiency, time, and automation, there is a potential selection bias toward studies framed around productivity gains
- relevant work reporting neutral or negative impacts — or focusing on other pedagogical outcomes without those terms — may be underrepresented."
Open questions raised
- Standardized and transparent reporting of outcome measures (including objective time-use metrics and validated quality rubrics)
- Stronger comparative designs (controlled field studies, quasi-experiments, and replication studies across institutions)
- Broader coverage of educational levels, regions, and resource-constrained settings
- Implementation research disentangling model-level effects from platform/workflow factors (prompting support, interface design, teacher training)
- More consistent evaluation of governance issues (privacy, academic integrity, bias mitigation, cost) to support responsible and scalable deployment
- Longitudinal, classroom-based studies in underrepresented contexts (pre-tertiary, rural, low-resource, multilingual settings)
Explore related topics
Related papers
- What Is the Impact of ChatGPT on Education? A Rapid Review of the LiteratureChung Kwan Lo · 2023 · 1,725 citations
- ChatGPT: Bullshit spewer or the end of traditional assessments in higher education?Jürgen Rudolph · 2023 · 1,674 citations
- Artificial intelligence in higher education: the state of the fieldHelen Crompton · 2023 · 1,378 citations
- Ethics of AI in Education: Towards a Community-Wide FrameworkW. Holmes · 2021 · 1,056 citations
- The effects of over-reliance on AI dialogue systems on students' cognitive abilities: a systematic reviewChunpeng Zhai · 2024 · 1,009 citations
- Shaping the Future of Education: Exploring the Potential and Consequences of AI and ChatGPT in Educational SettingsSimone Grassini · 2023 · 921 citations