12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Errors in AI-Assisted Retrieval of Medical Literature: A Comparative Study

Jenny Gao, Yongfeng Zhang, Mary L. Disis, Lanjing Zhang · arXiv (Cornell University) · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
2/4
Quality (LMQS)
E
Evidence
0
Citations
0.00
FWCI

Methodology & findings

Study design

Comparative observational study.

Sample

N = 2000, 9 groups

Primary method

Multivariable logistic regression analysis to examine independent associations of journal, publication year, and LLM platform with score ratio and complete miss rate. Chi-square or Fisher exact test used to compare categorical data by metrics. All statistical analyses performed using Stata (version 18, College Station, TX: StataCorp LLC). Statistical significance defined as p < 0.05.

Main result

The study found that "LLM platforms completely failed to retrieve correct reference data 47.8% of the time" with "the average score ratio of the 5 LLM platforms was 0.29 (standard deviation, 0.35; range, 0-1.25), with a higher score ratio indicating a higher accuracy in retrieving relevant references and correct bibliographic data." Additionally, "The highest and lowest accuracies were achieved by Grok (0.57) and Gemini (0.11), respectively."

Reports effect sizes and confidence intervals.

Research paradigm

Positivist/empiricist

Author conclusions

The authors concluded that "while LLMs hold promises as tools for biomedical literature exploration, their current performance in medical literature search and retrieval is highly variable and frequently unreliable." They further state that "Until robust hallucination mitigation, verification systems, and trustworthy AI frameworks are widely implemented, LLMs should be used as assistive, not authoritative, tools in biomedical scholarship."

Risk of bias

Selection bias: Only free-version LLM platforms tested; paid versions not evaluated; Selection bias: Limited to abstracts rather than full-text articles, potentially limiting retrieval performance; Detection bias: Manual validation of references involved human judgment in relevance scoring, introducing potential subjectivity; Confounding: Abstract length variability by journal (NEJM 250 words vs JAMA 350 words) may confound journal-specific performance differences; Temporal confounding: LLM models continue to improve; study conducted across July-August and December 2025, potentially introducing version differences; Measurement bias: Composite score ratio calculation with variable score caps may not equally weight all metrics; Selection bias: Only free versions of LLM platforms evaluated, not comprehensive/premium versions; Subjective bias: Human judgment involved in relevance scoring despite systematic approach; Limitation of scope: Only abstracts used, not full articles, which may constrain LLM performance; Temporal limitation: LLM capabilities improve rapidly; study findings may become outdated; Journal selection bias: Only four leading medical journals represented; results may not generalize to other journals; Selection bias: Only free-version LLM platforms evaluated, may not represent full capabilities of paid versions; Measurement bias: Relevance scoring involved human judgment which the authors acknowledge may be subjective and biased; Temporal bias: LLM capabilities improve rapidly, findings may be outdated quickly; Limited scope: Only abstracts used for LLM search, not full articles, may underestimate actual performance

Limitations

  • The authors stated that "We evaluated only free versions of LLM platforms and focused on a simple prompt design, which may not capture the full range of model capabilities." Additionally, "our relevance scoring, while systematic and mostly objective (e.g., citation by the original article), involved human judgment and may be subjective and biased." Furthermore, "only abstracts of the original articles were subject to the LLM search but may have limited its performance."

Open questions raised

  • The authors identify that: (1) More research is needed to understand and improve LLM-assisted literature retrieval; (2) Future works evaluating full articles (rather than abstracts only) would be interesting but require substantially more time; (3) More effective strategies for improving retrieval accuracy through alternative prompt designs are needed; (4) Investigation of paid/premium LLM versions compared to free versions is warranted.
  • Need for research on full articles rather than abstracts to improve LLM performance assessment
  • Development of robust hallucination mitigation strategies
  • Implementation of better verification systems and trustworthy AI frameworks
  • Further investigation of factors causing variable performance across LLM platforms and journals
  • Evaluation of premium/paid versions of LLM platforms beyond free versions
Data: The authors state: "The data is available from the corresponding authors on reasonable request." No public dataset repository URL is provided.; Data available from corresponding authors on reasonable request (as stated: "The data is available from the corresponding authors on reasonable request."); "The data is available from the corresponding authors on reasonable request."Extracted from: pdfAgreement 52%

Explore related topics

Related papers