12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Strategies for developing AI competencies in higher education

Miroslava Nadkova Petrova · Frontiers in Education · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

7/10
Relevance
0/4
Quality (LMQS)
C
Evidence
4
Citations
25.96
FWCI
Top 10%
Impact

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.3389/feduc.2025.1683909

Methodology & findings

Study design

LLM-based Delphi simulation conducted in four iterative phases.

Sample

N = 20, 5 groups

Primary method

Median Likert scores calculated from 5-point scale ratings (1-5); interquartile range (IQR) reported for variability; consensus threshold defined as median score ≥4.0 for validated themes and <3.5 for disputed themes; inductive thematic analysis on qualitative responses; clustering-based comparison for consensus analysis; two-stage sensitivity analysis comparing outputs across multiple prompt strategies and LLM models (DeepSeek-V3 vs. GPT-5); descriptive statistics for expert panel demographics.

Main result

The study found that "the research process employed large language models (LLMs) to conduct a simulated exploration with inductive thematic analysis of interdisciplinary perspectives, prioritize critical themes through iterative rating cycles, and resolve polarization via structured deliberation of disputed concepts." Key outputs include "the development of a consensus framework outlining universal AI literacy standards, human-AI collaborative pedagogy models, equity-centered implementation protocols, and ethical guardrails for responsible adoption, together with a toolkit with practical guidelines to operationalize the consensus findings." Survey results showed that experts broadly endorsed the four themes of the simulated consensus, achieving "strong agreement across the six OECD evaluation criteria with median Likert score ≥4.0, with the sole exception of ethical guardrails' sustainability, which had a median of 3.5."

Reports effect sizes and confidence intervals.

Research paradigm

Interpretivist with computational simulation elements

Author conclusions

The author concludes that "the study aims to assess AI's potential as a collaborative agent in educational design and to evaluate to what extent an AI-generated framework meets established OECD criteria for quality and robustness." The research demonstrates that an LLM-based Delphi methodology can generate actionable frameworks for AI competency development, with the validation results showing strong expert agreement (median Likert ≥4.0 across most dimensions) on the proposed consensus framework and implementation toolkit. The author notes that "experts broadly endorsed the four themes of the simulated consensus, achieving strong agreement across the six OECD evaluation criteria with median Likert score ≥4.0, with the sole exception of ethical guardrails' sustainability, which had a median of 3.5."

Risk of bias

Researcher bias in prompt design and synthesis; LLM hallucination and knowledge cutoff limitations (DeepSeek-V3 knowledge cutoff: June 2025); Simulation bias: LLM personas may not accurately represent diverse expert perspectives; Small validation panel (n=8) limiting representativeness; Single primary LLM with limited comparator testing; Potential temperature and parameter selection bias (temperature set to 1.0, top_p=0.9); No human expert participation in the initial Delphi rounds, only in validation; Simulation bias: LLM-generated expert perspectives may not accurately represent real expert consensus; Training data bias: DeepSeek-V3 knowledge cutoff (June 2025) may introduce temporal bias; Model parameter selection: Default temperature (1.0), top_p (0.9) may influence consistency of outputs; Small validation panel: Only 8 human experts for validation reduces generalizability; Lack of actual expert participation: No human experts participated in the Delphi rounds themselves; Interface non-determinism: Acknowledged mitigation via multiple runs but not fully eliminated; LLM hallucination and biases in training data (DeepSeek-V3); Lack of actual human expert participation in initial Delphi rounds; Selection bias in validation panel (N=8, recruited through targeted invitations); Interface non-determinism mitigated through median outcomes but not fully eliminated; Single-run versus multi-run output variability; Potential gender imbalance in validation panel (62.5% male)

Limitations

  • The study is limited by the fact that "the researcher controlled the entire process, designing and refining the prompts, synthesizing the outputs and conducting the final analysis," which may introduce researcher bias
  • Additionally, the simulated expert panel lacks the lived experience and tacit knowledge of actual human experts
  • The paper notes that "while the traditional method relies on human expertise, recent advances in AI, particularly the breakthroughs in large language models (LLMs) capable of complex cross-domain reasoning, now enable AI to replicate human cognition in this process," but the degree to which LLM-simulated expertise matches genuine expert judgment remains unclear
  • Furthermore, the validation expert panel was small (n=8) and the study was conducted using a single primary LLM with comparison testing against only one alternative model.

Open questions raised

  • The authors identify the following gaps: (1) 'Relatively small number of institutions are actively engaged in evaluation measures for AI's impact, which signals significant gaps in comprehensive policy development, communication strategies, and equitable distribution of resources for generative AI integration.' (2) Areas for further study include pilot research on AI detection tools, holographic teaching effectiveness, and the safety/feasibility of mental health automation with clinician monitoring. (3) The need for 'new ways of assessing learning' as many current methods (essays, summaries, reports) are vulnerable to AI. (4) Limited research on human-AI collaborative capabilities in authentic contexts.
  • Need for empirical research on the effectiveness of AI-enhanced pedagogical methods in actual classroom settings
  • Assessment validity: whether AI-augmented creative work can be evaluated fairly and long-term developmental effects on student creativity
  • Piloting of AI detection tools to address potential bias against non-native speakers and effectiveness of explainable AI
  • Research on holographic teaching impact on actual learning comprehension and skill acquisition
  • Clinician-monitored pilot studies on AI mental health automation to test feasibility and safety
Extracted from: pdfAgreement 54%

Explore related topics

Related papers