Peer-reviewed by human experts: AI failed in key steps to generate a scoping review on the neural mechanisms of cross-education
Marco Morrone, Tibor Hortobágyi, Dianna Kidgell, J. P. Farthing, Franca Deriu, Andrea Manca · European Journal of Applied Physiology · 2025
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1007/s00421-025-06100-w
Methodology & findings
Study design
Critical peer-review appraisal methodology.
Sample
N = 4, 4 groups
Primary method
Consensus-based qualitative synthesis of expert reviewer feedback. No quantitative statistical analyses performed. Authors state: 'After consultation among the article curators, major issues from each expert were identified, extracted and attributed to experts by consensus.' Feedback was documented and organized into structured table (Table 1) summarizing reviewer assessments by article section.
Main result
The study found that the LLM-generated scoping review "did not perform key steps of the framework that is commonly adopted for conducting a scoping review (Arksey and O'Malley 2005), i.e., identifying the research questions, selecting studies to be included in the review comprehensively, charting the data, collating, summarizing and reporting the results of such a process. Instead, it proved to be a narrative, non-systematic poorly structured review." Key specific failures included: "84% (52/62) were from open-access sources" with systematic exclusion of subscription-based foundational literature; "Countless inaccuracies and errors in the background and, overall, in the whole manuscript" regarding references; and "The LLM is below-the-bar in managing scientific bibliography."
Reports effect sizes.
Research paradigm
Critical/Evaluative - Meta-scientific analysis of AI capabilities in academic writing
Author conclusions
"The main finding of this analysis is that this LLM-generated scoping review did not perform key steps of the framework that is commonly adopted for conducting a scoping review (Arksey and O'Malley 2005), i.e., identifying the research questions, selecting studies to be included in the review comprehensively, charting the data, collating, summarizing and reporting the results of such a process. Instead, it proved to be a narrative, non-systematic poorly structured review. This conclusion doesn't mean that scholars should give up on the use of LLMs for scientific writing." They conclude that "even though LLMs seem still unable to replace subject matter experts, they promise to transform how reviews are prepared, 'shifting the human role from exhaustive curator to creative synthesizer, empowered by' [context windows and capabilities yet to be fully realized]."
Risk of bias
Interpretive bias: Authors served as both investigators and subject-matter experts in cross-education research; Selection bias: Four expert reviewers were necessarily specialists in the specific field being reviewed; Single LLM tested: Only Gemini 2.5 Pro evaluated; no comparison with other LLMs; Single-shot prompting: No iterative refinement based on initial results; Narrow domain: Focused on healthy individuals' cross-education, limiting generalizability; Training data cutoff: LLM had September 2024 cutoff, missing recent studies by Lecce et al. (2025a, b); Selection bias in LLM output: Only 1 of 3 generations was complete; remaining two were similar, creating potential selection bias; Interpretive bias: Authors were both investigators and domain experts in the subject matter; Access bias: LLM systematically favored open-access literature (84% of citations) while excluding subscription-based foundational works; Referencing bias: LLM relied on paywalled sources through ResearchGate and other platforms, potentially leading to copyright violations; Data currency bias: LLM used arbitrary cut-off date of September 2024 (7 months before generation), missing recent works; Prompt design bias: Single prompt iteration without refinement based on results; Interpretive bias: authors served dual role as both investigators and domain experts in cross-education neurophysiology; Selection bias in AI output: only one of three LLM generations was selected for review (first complete one); other iterations not analyzed; Technology-specific constraints: LLM training data cutoff excluded recent studies (e.g., Lecce et al. 2025a,b); paywalled journal access limitations created systematic bias toward open-access sources (84% of citations); Prompt design effect: single-shot prompting strategy (not iteratively refined) may not reflect best-case LLM performance; Institutional access bias: suggests that "running LLM inside or outside of an academic institution would make a difference in the final LLM-script, depending on the portfolio of publishers and journals subscribed by the scholar's institution"; Reviewer expertise bias: all four peer reviewers were leading experts in the specific field of cross-education, potentially biasing assessment toward domain-specific standards
Limitations
- "A key limitation of this study is its intentionally narrow focus on cross-education mechanisms in healthy individuals, the authors' main research area
- While this provided a controlled framework to test the LLM's ability to generate a scientific review, it also restricted the study's scope and generalizability to pathological conditions
- We also acknowledge the dual role of the authors as both investigators and subject-matter experts
- While expertise in this domain was critical for identifying subtle factual errors and hallucinations that a generalist might miss, it introduces a potential interpretive bias." Additionally, "the scope of our project, which was to evaluate the Deep Research out-of-the-box capabilities, rather than best-case performance achievable through expert prompt engineering, multi-agent frameworks, or fine-tuning."
Open questions raised
- Need for credible Reference Manager systems that can access both open-access and subscription-based databases
- Development of mechanisms to assess quality of evidence and hierarchize findings within LLM outputs
- Incorporation of evidence-based frameworks (GRADE methodology) into LLM fine-tuning
- Multi-agent workflows and Retrieval-Augmented Generation with structured metadata
- Formal agreements between AI developers and academic publishers
- Testing across multiple LLMs and comparison of capabilities
Explore related topics
Related papers
- PRISMA Extension for Scoping Reviews (PRISMA-ScR): Checklist and ExplanationAndrea C. Tricco · 2018 · 40,391 citations
- Rayyan—a web and mobile app for systematic reviewsMourad Ouzzani · 2016 · 24,664 citations
- Cochrane Handbook for Systematic Reviews of Interventions2019 · 14,420 citations
- Guidance for conducting systematic scoping reviewsMicah D.J. Peters · 2015 · 7,472 citations
- Updated methodological guidance for the conduct of scoping reviewsMicah D.J. Peters · 2020 · 6,688 citations
- Which academic search systems are suitable for systematic reviews or meta‐analyses? Evaluating retrieval qualities of Google Scholar, PubMed, and 26 other resourcesMichael Gusenbauer · 2019 · 2,116 citations