Errors in AI-Assisted Retrieval of Medical Literature: A Comparative Study
Jenny Gao, Yongfeng Zhang, Mary L. Disis, Lanjing Zhang · arXiv (Cornell University) · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
Methodology & findings
Study design
Comparative observational study.
Sample
N = 2000, 9 groups
Primary method
Multivariable logistic regression analysis to examine independent associations of journal, publication year, and LLM platform with score ratio and complete miss rate. Chi-square or Fisher exact test used to compare categorical data by metrics. All statistical analyses performed using Stata (version 18, College Station, TX: StataCorp LLC). Statistical significance defined as p < 0.05.
Main result
The study found that "LLM platforms completely failed to retrieve correct reference data 47.8% of the time" with "the average score ratio of the 5 LLM platforms was 0.29 (standard deviation, 0.35; range, 0-1.25), with a higher score ratio indicating a higher accuracy in retrieving relevant references and correct bibliographic data." Additionally, "The highest and lowest accuracies were achieved by Grok (0.57) and Gemini (0.11), respectively."
Reports effect sizes and confidence intervals.
Research paradigm
Positivist/empiricist
Author conclusions
The authors concluded that "while LLMs hold promises as tools for biomedical literature exploration, their current performance in medical literature search and retrieval is highly variable and frequently unreliable." They further state that "Until robust hallucination mitigation, verification systems, and trustworthy AI frameworks are widely implemented, LLMs should be used as assistive, not authoritative, tools in biomedical scholarship."
Risk of bias
Selection bias: Only free-version LLM platforms tested; paid versions not evaluated; Selection bias: Limited to abstracts rather than full-text articles, potentially limiting retrieval performance; Detection bias: Manual validation of references involved human judgment in relevance scoring, introducing potential subjectivity; Confounding: Abstract length variability by journal (NEJM 250 words vs JAMA 350 words) may confound journal-specific performance differences; Temporal confounding: LLM models continue to improve; study conducted across July-August and December 2025, potentially introducing version differences; Measurement bias: Composite score ratio calculation with variable score caps may not equally weight all metrics; Selection bias: Only free versions of LLM platforms evaluated, not comprehensive/premium versions; Subjective bias: Human judgment involved in relevance scoring despite systematic approach; Limitation of scope: Only abstracts used, not full articles, which may constrain LLM performance; Temporal limitation: LLM capabilities improve rapidly; study findings may become outdated; Journal selection bias: Only four leading medical journals represented; results may not generalize to other journals; Selection bias: Only free-version LLM platforms evaluated, may not represent full capabilities of paid versions; Measurement bias: Relevance scoring involved human judgment which the authors acknowledge may be subjective and biased; Temporal bias: LLM capabilities improve rapidly, findings may be outdated quickly; Limited scope: Only abstracts used for LLM search, not full articles, may underestimate actual performance
Limitations
- The authors stated that "We evaluated only free versions of LLM platforms and focused on a simple prompt design, which may not capture the full range of model capabilities." Additionally, "our relevance scoring, while systematic and mostly objective (e.g., citation by the original article), involved human judgment and may be subjective and biased." Furthermore, "only abstracts of the original articles were subject to the LLM search but may have limited its performance."
Open questions raised
- The authors identify that: (1) More research is needed to understand and improve LLM-assisted literature retrieval; (2) Future works evaluating full articles (rather than abstracts only) would be interesting but require substantially more time; (3) More effective strategies for improving retrieval accuracy through alternative prompt designs are needed; (4) Investigation of paid/premium LLM versions compared to free versions is warranted.
- Need for research on full articles rather than abstracts to improve LLM performance assessment
- Development of robust hallucination mitigation strategies
- Implementation of better verification systems and trustworthy AI frameworks
- Further investigation of factors causing variable performance across LLM platforms and journals
- Evaluation of premium/paid versions of LLM platforms beyond free versions
Explore related topics
Related papers
- Rayyan—a web and mobile app for systematic reviewsMourad Ouzzani · 2016 · 24,664 citations
- Guidance for conducting systematic scoping reviewsMicah D.J. Peters · 2015 · 7,472 citations
- Which academic search systems are suitable for systematic reviews or meta‐analyses? Evaluating retrieval qualities of Google Scholar, PubMed, and 26 other resourcesMichael Gusenbauer · 2019 · 2,116 citations
- Comparing scientific abstracts generated by ChatGPT to real abstracts with detectors and blinded human reviewersCatherine A. Gao · 2023 · 657 citations
- Fabrication and errors in the bibliographic citations generated by ChatGPTWilliam H. Walters · 2023 · 352 citations
- Human-AI collaboration patterns in AI-assisted academic writingAndy Nguyen · 2024 · 301 citations