12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Peer-reviewed by human experts: AI failed in key steps to generate a scoping review on the neural mechanisms of cross-education

Marco Morrone, Tibor Hortobágyi, Dianna Kidgell, J. P. Farthing, Franca Deriu, Andrea Manca · European Journal of Applied Physiology · 2025

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

10/10
Relevance
2/4
Quality (LMQS)
E
Evidence
1
Citations
0.44
FWCI
Top 10%
Impact

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1007/s00421-025-06100-w

Methodology & findings

Study design

Critical peer-review appraisal methodology.

Sample

N = 4, 4 groups

Primary method

Consensus-based qualitative synthesis of expert reviewer feedback. No quantitative statistical analyses performed. Authors state: 'After consultation among the article curators, major issues from each expert were identified, extracted and attributed to experts by consensus.' Feedback was documented and organized into structured table (Table 1) summarizing reviewer assessments by article section.

Main result

The study found that the LLM-generated scoping review "did not perform key steps of the framework that is commonly adopted for conducting a scoping review (Arksey and O'Malley 2005), i.e., identifying the research questions, selecting studies to be included in the review comprehensively, charting the data, collating, summarizing and reporting the results of such a process. Instead, it proved to be a narrative, non-systematic poorly structured review." Key specific failures included: "84% (52/62) were from open-access sources" with systematic exclusion of subscription-based foundational literature; "Countless inaccuracies and errors in the background and, overall, in the whole manuscript" regarding references; and "The LLM is below-the-bar in managing scientific bibliography."

Reports effect sizes.

Research paradigm

Critical/Evaluative - Meta-scientific analysis of AI capabilities in academic writing

Author conclusions

"The main finding of this analysis is that this LLM-generated scoping review did not perform key steps of the framework that is commonly adopted for conducting a scoping review (Arksey and O'Malley 2005), i.e., identifying the research questions, selecting studies to be included in the review comprehensively, charting the data, collating, summarizing and reporting the results of such a process. Instead, it proved to be a narrative, non-systematic poorly structured review. This conclusion doesn't mean that scholars should give up on the use of LLMs for scientific writing." They conclude that "even though LLMs seem still unable to replace subject matter experts, they promise to transform how reviews are prepared, 'shifting the human role from exhaustive curator to creative synthesizer, empowered by' [context windows and capabilities yet to be fully realized]."

Risk of bias

Interpretive bias: Authors served as both investigators and subject-matter experts in cross-education research; Selection bias: Four expert reviewers were necessarily specialists in the specific field being reviewed; Single LLM tested: Only Gemini 2.5 Pro evaluated; no comparison with other LLMs; Single-shot prompting: No iterative refinement based on initial results; Narrow domain: Focused on healthy individuals' cross-education, limiting generalizability; Training data cutoff: LLM had September 2024 cutoff, missing recent studies by Lecce et al. (2025a, b); Selection bias in LLM output: Only 1 of 3 generations was complete; remaining two were similar, creating potential selection bias; Interpretive bias: Authors were both investigators and domain experts in the subject matter; Access bias: LLM systematically favored open-access literature (84% of citations) while excluding subscription-based foundational works; Referencing bias: LLM relied on paywalled sources through ResearchGate and other platforms, potentially leading to copyright violations; Data currency bias: LLM used arbitrary cut-off date of September 2024 (7 months before generation), missing recent works; Prompt design bias: Single prompt iteration without refinement based on results; Interpretive bias: authors served dual role as both investigators and domain experts in cross-education neurophysiology; Selection bias in AI output: only one of three LLM generations was selected for review (first complete one); other iterations not analyzed; Technology-specific constraints: LLM training data cutoff excluded recent studies (e.g., Lecce et al. 2025a,b); paywalled journal access limitations created systematic bias toward open-access sources (84% of citations); Prompt design effect: single-shot prompting strategy (not iteratively refined) may not reflect best-case LLM performance; Institutional access bias: suggests that "running LLM inside or outside of an academic institution would make a difference in the final LLM-script, depending on the portfolio of publishers and journals subscribed by the scholar's institution"; Reviewer expertise bias: all four peer reviewers were leading experts in the specific field of cross-education, potentially biasing assessment toward domain-specific standards

Limitations

  • "A key limitation of this study is its intentionally narrow focus on cross-education mechanisms in healthy individuals, the authors' main research area
  • While this provided a controlled framework to test the LLM's ability to generate a scientific review, it also restricted the study's scope and generalizability to pathological conditions
  • We also acknowledge the dual role of the authors as both investigators and subject-matter experts
  • While expertise in this domain was critical for identifying subtle factual errors and hallucinations that a generalist might miss, it introduces a potential interpretive bias." Additionally, "the scope of our project, which was to evaluate the Deep Research out-of-the-box capabilities, rather than best-case performance achievable through expert prompt engineering, multi-agent frameworks, or fine-tuning."

Open questions raised

  • Need for credible Reference Manager systems that can access both open-access and subscription-based databases
  • Development of mechanisms to assess quality of evidence and hierarchize findings within LLM outputs
  • Incorporation of evidence-based frameworks (GRADE methodology) into LLM fine-tuning
  • Multi-agent workflows and Retrieval-Augmented Generation with structured metadata
  • Formal agreements between AI developers and academic publishers
  • Testing across multiple LLMs and comparison of capabilities
Extracted from: pdfAgreement 58%

Explore related topics

Related papers