12,637 papers · updated 18 Sept 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Hallucination Rates and Reference Accuracy of ChatGPT and Bard for Systematic Reviews: Comparative Analysis

Mikaël Chelli, Jules Descamps, Vincent Lavoué, Christophe Trojani, Michel Azar, Marcel Deckert et al. · Journal of Medical Internet Research · 2024

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
E
Evidence
295
Citations
30.17
FWCI
Top 10%
Impact

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.2196/53164

Methodology & findings

Study design

Comparative empirical evaluation using sequential design.

Sample

N = 471, 4 groups

Primary method

Chi-square tests to compare performance metrics and categorical variables (authors' nationalities, open-access status) between LLMs. Significance threshold P<.05. Metrics calculated: recall (TP/(TP+FN)), precision (TP/(TP+FP)), F1-score (2×precision×recall/(precision+recall)). Hallucination rate calculated as proportion of LLM-generated references with ≥2 incorrect data points among title, first author, or year. Statistical analysis performed with EasyMedStat (version 3.24).

Main result

The study found that "Precision rates for GPT-3.5, GPT-4, and Bard were 9.4% (13/139), 13.4% (16/119), and 0% (0/104) respectively (P<.001). Recall rates were 11.9% (13/109) for GPT-3.5 and 13.7% (15/109) for GPT-4, with Bard failing to retrieve any relevant papers (P<.001). Hallucination rates stood at 39.6% (55/139) for GPT-3.5, 28.6% (34/119) for GPT-4, and 91.4% (95/104) for Bard (P<.001)." The most important finding is that "using LLMs such as ChatGPT and Bard to conduct systematic reviews for a common condition such as rotator cuff disease can generate misleading or 'hallucinated' references, exceeding a 25% rate."

Reports effect sizes.

Research paradigm

Empirical-positivist (comparative measurement of LLM performance against gold-standard systematic review references)

Author conclusions

"ChatGPT and Bard exhibit the capacity to generate convincingly authentic references for systematic reviews but also yield hallucinated papers in 28.6% (34/119) to 91.3% (95/104) of cases. Among the models tested, GPT-4 displayed superior performance in generating legitimate and relevant references but, like the other models, largely failed to respect the established eligibility criteria. Given their current state, LLMs such as ChatGPT and Bard should not be used as the sole or primary means for conducting systematic reviews of literature, and it is crucial that references generated by these tools undergo rigorous validation by the authors of scientific papers."

Risk of bias

Selection bias: Study restricted to shoulder rotator cuff pathology; generalizability to other medical fields unknown; Model selection bias: Only 3 LLMs tested; other models may perform differently; Prompt engineering bias: Study tested 2 prompt versions but acknowledged alternative prompts could yield different results; Gold standard bias: Original systematic reviews serve as ground truth, but may themselves contain errors or omissions; Publication bias: Study focuses on published papers only; grey literature not included in assessment; Geographic bias: LLMs demonstrated preference for American authors (44% vs 16.5% in original reviews, P<.001); Open-access bias: LLMs preferentially retrieved open-access papers (38% vs 27.5% in originals); Temporal bias: Search restricted to 2020 publications, limiting applicability to current LLM capabilities; landscape continually evolving; Prompt bias: Authors acknowledge that choice of prompt plays crucial role; findings may not generalize to alternative query formulations; Publication bias: LLMs may have biases from training data; Confirmation bias: Authors had clinical expertise in rotator cuff pathology, potentially influencing interpretations; Prompt bias: Choice of prompt formulation may influence results

Limitations

  • "The scope of the study was exclusively focused on systematic reviews related to shoulder rotator cuff pathology
  • Consequently, it must be recognized that the findings might not be universally applicable across diverse medical specialties or disciplines." Additionally, "The examination was also restricted to 3 LLMs, specifically GPT-3.5, GPT-4, and Bard" and "the field lacks established guidelines for leveraging LLMs to optimize accuracy
  • Notwithstanding rigorous attempts to devise specific, comprehensive prompts, it remains plausible that alternative queries could generate more precise outcomes." The authors also note "Our decision not to provide the initial PubMed results list to LLMs for assessing paper eligibility was deliberate, aimed at preserving study integrity and interpretability
  • While providing the list might enhance LLM accuracy, it introduces bias by guiding models toward replicating the provided set rather than autonomously identifying relevant studies."

Open questions raised

  • Need for established guidelines on optimizing LLM prompts for academic tasks
  • further investigation needed across diverse medical fields to assess definitively whether LLMs introduce geographic and demographic biases
  • necessity for refining LLM training and functionality before deployment in rigorous academic purposes
  • need for scholarly usage guidelines integrated into LLM software itself to outline lack of liability for citation inaccuracies
  • The authors identify needs for:
  • established guidelines for leveraging LLMs to optimize accuracy
Data: "The datasets generated and analyzed during this study are available from the corresponding author on reasonable request." No public repository or URL provided.Code: None mentionedExtracted from: pdf

Explore related topics

Related papers