Strategies for developing AI competencies in higher education
Miroslava Nadkova Petrova · Frontiers in Education · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.3389/feduc.2025.1683909
Methodology & findings
Study design
LLM-based Delphi simulation conducted in four iterative phases.
Sample
N = 20, 5 groups
Primary method
Median Likert scores calculated from 5-point scale ratings (1-5); interquartile range (IQR) reported for variability; consensus threshold defined as median score ≥4.0 for validated themes and <3.5 for disputed themes; inductive thematic analysis on qualitative responses; clustering-based comparison for consensus analysis; two-stage sensitivity analysis comparing outputs across multiple prompt strategies and LLM models (DeepSeek-V3 vs. GPT-5); descriptive statistics for expert panel demographics.
Main result
The study found that "the research process employed large language models (LLMs) to conduct a simulated exploration with inductive thematic analysis of interdisciplinary perspectives, prioritize critical themes through iterative rating cycles, and resolve polarization via structured deliberation of disputed concepts." Key outputs include "the development of a consensus framework outlining universal AI literacy standards, human-AI collaborative pedagogy models, equity-centered implementation protocols, and ethical guardrails for responsible adoption, together with a toolkit with practical guidelines to operationalize the consensus findings." Survey results showed that experts broadly endorsed the four themes of the simulated consensus, achieving "strong agreement across the six OECD evaluation criteria with median Likert score ≥4.0, with the sole exception of ethical guardrails' sustainability, which had a median of 3.5."
Reports effect sizes and confidence intervals.
Research paradigm
Interpretivist with computational simulation elements
Author conclusions
The author concludes that "the study aims to assess AI's potential as a collaborative agent in educational design and to evaluate to what extent an AI-generated framework meets established OECD criteria for quality and robustness." The research demonstrates that an LLM-based Delphi methodology can generate actionable frameworks for AI competency development, with the validation results showing strong expert agreement (median Likert ≥4.0 across most dimensions) on the proposed consensus framework and implementation toolkit. The author notes that "experts broadly endorsed the four themes of the simulated consensus, achieving strong agreement across the six OECD evaluation criteria with median Likert score ≥4.0, with the sole exception of ethical guardrails' sustainability, which had a median of 3.5."
Risk of bias
Researcher bias in prompt design and synthesis; LLM hallucination and knowledge cutoff limitations (DeepSeek-V3 knowledge cutoff: June 2025); Simulation bias: LLM personas may not accurately represent diverse expert perspectives; Small validation panel (n=8) limiting representativeness; Single primary LLM with limited comparator testing; Potential temperature and parameter selection bias (temperature set to 1.0, top_p=0.9); No human expert participation in the initial Delphi rounds, only in validation; Simulation bias: LLM-generated expert perspectives may not accurately represent real expert consensus; Training data bias: DeepSeek-V3 knowledge cutoff (June 2025) may introduce temporal bias; Model parameter selection: Default temperature (1.0), top_p (0.9) may influence consistency of outputs; Small validation panel: Only 8 human experts for validation reduces generalizability; Lack of actual expert participation: No human experts participated in the Delphi rounds themselves; Interface non-determinism: Acknowledged mitigation via multiple runs but not fully eliminated; LLM hallucination and biases in training data (DeepSeek-V3); Lack of actual human expert participation in initial Delphi rounds; Selection bias in validation panel (N=8, recruited through targeted invitations); Interface non-determinism mitigated through median outcomes but not fully eliminated; Single-run versus multi-run output variability; Potential gender imbalance in validation panel (62.5% male)
Limitations
- The study is limited by the fact that "the researcher controlled the entire process, designing and refining the prompts, synthesizing the outputs and conducting the final analysis," which may introduce researcher bias
- Additionally, the simulated expert panel lacks the lived experience and tacit knowledge of actual human experts
- The paper notes that "while the traditional method relies on human expertise, recent advances in AI, particularly the breakthroughs in large language models (LLMs) capable of complex cross-domain reasoning, now enable AI to replicate human cognition in this process," but the degree to which LLM-simulated expertise matches genuine expert judgment remains unclear
- Furthermore, the validation expert panel was small (n=8) and the study was conducted using a single primary LLM with comparison testing against only one alternative model.
Open questions raised
- The authors identify the following gaps: (1) 'Relatively small number of institutions are actively engaged in evaluation measures for AI's impact, which signals significant gaps in comprehensive policy development, communication strategies, and equitable distribution of resources for generative AI integration.' (2) Areas for further study include pilot research on AI detection tools, holographic teaching effectiveness, and the safety/feasibility of mental health automation with clinician monitoring. (3) The need for 'new ways of assessing learning' as many current methods (essays, summaries, reports) are vulnerable to AI. (4) Limited research on human-AI collaborative capabilities in authentic contexts.
- Need for empirical research on the effectiveness of AI-enhanced pedagogical methods in actual classroom settings
- Assessment validity: whether AI-augmented creative work can be evaluated fairly and long-term developmental effects on student creativity
- Piloting of AI detection tools to address potential bias against non-native speakers and effectiveness of explainable AI
- Research on holographic teaching impact on actual learning comprehension and skill acquisition
- Clinician-monitored pilot studies on AI mental health automation to test feasibility and safety
Explore related topics
Related papers
- What if the devil is my guardian angel: ChatGPT as a case study of using chatbots in educationAhmed Tlili · 2023 · 1,587 citations
- Conceptualizing AI literacy: An exploratory reviewDavy Tsz Kit Ng · 2021 · 1,492 citations
- A SWOT analysis of ChatGPT: Implications for educational practice and researchMohammadreza Farrokhnia · 2023 · 1,171 citations
- A comprehensive AI policy education framework for university teaching and learningCecilia Ka Yuk Chan · 2023 · 1,160 citations
- Shaping the Future of Education: Exploring the Potential and Consequences of AI and ChatGPT in Educational SettingsSimone Grassini · 2023 · 921 citations
- ChatGPT for Education and Research: Opportunities, Threats, and StrategiesMd. Mostafizer Rahman · 2023 · 904 citations