12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Hallucination Rates and Reference Accuracy of ChatGPT and Bard for Systematic Reviews: Comparative Analysis

Mikaël Chelli, Jules Descamps, Vincent Lavoué, Christophe Trojani, Michel Azar, Marcel Deckert et al. · Journal of Medical Internet Research · 2024

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
3/4
Quality (LMQS)
E
Evidence
295
Citations
30.17
FWCI
Top 10%
Impact

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.2196/53164

Methodology & findings

Study design

Comparative empirical evaluation using sequential design.

Sample

N = 471, 10 groups

Primary method

Chi-square tests to compare performance metrics and categorical variables (authors' nationalities, open-access status) between LLMs. Significance threshold P<.05. Metrics calculated: recall (TP/(TP+FN)), precision (TP/(TP+FP)), F1-score (2×precision×recall/(precision+recall)). Hallucination rate calculated as proportion of LLM-generated references with ≥2 incorrect data points among title, first author, or year. Statistical analysis performed with EasyMedStat (version 3.24).

Main result

The study found that "Precision rates for GPT-3.5, GPT-4, and Bard were 9.4% (13/139), 13.4% (16/119), and 0% (0/104) respectively (P<.001). Recall rates were 11.9% (13/109) for GPT-3.5 and 13.7% (15/109) for GPT-4, with Bard failing to retrieve any relevant papers (P<.001). Hallucination rates stood at 39.6% (55/139) for GPT-3.5, 28.6% (34/119) for GPT-4, and 91.4% (95/104) for Bard (P<.001)." The most important finding is that "using LLMs such as ChatGPT and Bard to conduct systematic reviews for a common condition such as rotator cuff disease can generate misleading or 'hallucinated' references, exceeding a 25% rate."

Reports effect sizes.

Research paradigm

Empirical-positivist (comparative measurement of LLM performance against gold-standard systematic review references)

Author conclusions

"ChatGPT and Bard exhibit the capacity to generate convincingly authentic references for systematic reviews but also yield hallucinated papers in 28.6% (34/119) to 91.3% (95/104) of cases. Among the models tested, GPT-4 displayed superior performance in generating legitimate and relevant references but, like the other models, largely failed to respect the established eligibility criteria. Given their current state, LLMs such as ChatGPT and Bard should not be used as the sole or primary means for conducting systematic reviews of literature, and it is crucial that references generated by these tools undergo rigorous validation by the authors of scientific papers."

Risk of bias

Selection bias: Study restricted to shoulder rotator cuff pathology; generalizability to other medical fields unknown; Model selection bias: Only 3 LLMs tested; other models may perform differently; Prompt engineering bias: Study tested 2 prompt versions but acknowledged alternative prompts could yield different results; Gold standard bias: Original systematic reviews serve as ground truth, but may themselves contain errors or omissions; Publication bias: Study focuses on published papers only; grey literature not included in assessment; Geographic bias: LLMs demonstrated preference for American authors (44% vs 16.5% in original reviews, P<.001); Open-access bias: LLMs preferentially retrieved open-access papers (38% vs 27.5% in originals); Temporal bias: Search restricted to 2020 publications, limiting applicability to current LLM capabilities; Selection bias: Study focused exclusively on shoulder rotator cuff pathology, limiting generalizability; Model selection bias: Only 3 LLMs tested; landscape continually evolving; Prompt bias: Authors acknowledge that choice of prompt plays crucial role; findings may not generalize to alternative query formulations; Publication bias: LLMs may have biases from training data; study found overrepresentation of American authors (44% in GPT-3.5 vs 16.5% in original reviews); Confirmation bias: Authors had clinical expertise in rotator cuff pathology, potentially influencing interpretations; Selection bias: Only shoulder rotator cuff pathology reviews were included; Model-specific bias: Only 3 LLM models tested; Prompt bias: Choice of prompt formulation may influence results; Geographical bias: LLMs over-represented American authors (44% vs 16.5% in original reviews); Open-access bias: LLMs preferentially selected open-access papers; Training data bias: LLMs trained on biased datasets may perpetuate stereotypes

Limitations

  • "The scope of the study was exclusively focused on systematic reviews related to shoulder rotator cuff pathology
  • Consequently, it must be recognized that the findings might not be universally applicable across diverse medical specialties or disciplines." Additionally, "The examination was also restricted to 3 LLMs, specifically GPT-3.5, GPT-4, and Bard" and "the field lacks established guidelines for leveraging LLMs to optimize accuracy
  • Notwithstanding rigorous attempts to devise specific, comprehensive prompts, it remains plausible that alternative queries could generate more precise outcomes." The authors also note "Our decision not to provide the initial PubMed results list to LLMs for assessing paper eligibility was deliberate, aimed at preserving study integrity and interpretability
  • While providing the list might enhance LLM accuracy, it introduces bias by guiding models toward replicating the provided set rather than autonomously identifying relevant studies."

Open questions raised

  • Need for established guidelines on optimizing LLM prompts for academic tasks; further investigation needed across diverse medical fields to assess definitively whether LLMs introduce geographic and demographic biases; necessity for refining LLM training and functionality before deployment in rigorous academic purposes; need for scholarly usage guidelines integrated into LLM software itself to outline lack of liability for citation inaccuracies
  • The authors identify needs for: (1) further investigation across diverse medical fields to ascertain whether LLMs introduce geographical biases definitively; (2) established guidelines for leveraging LLMs to optimize accuracy; (3) refinement of LLM training and functionality before confident use for rigorous academic purposes; (4) prominent warning statements or scholarly usage guidelines integrated into LLM software itself before deployment.
  • The authors identify the need for "further investigation across diverse medical fields is warranted to ascertain whether these LLMs may introduce such biases definitively." They also note the field "lacks established guidelines for leveraging LLMs to optimize accuracy" and call for future work on optimizing prompts and developing standards for LLM use in systematic reviews.
Data: "The datasets generated and analyzed during this study are available from the corresponding author on reasonable request." No public repository or URL provided.; "The data sets generated and analyzed during this study are available from the corresponding author on reasonable request." Authors note: "All papers accessed by the large language models (LLMs) were publicly available, and no proprietary or subscription-based sources were used without appropriate access rights."; "All papers accessed by the large language models (LLMs) were publicly available, and no proprietary or subscription-based sources were used without appropriate access rights... The data sets generated and analyzed during this study are available from the corresponding author on reasonable request."Code: None mentionedExtracted from: pdfAgreement 54%

Explore related topics

Related papers