Hallucination Rates and Reference Accuracy of ChatGPT and Bard for Systematic Reviews: Comparative Analysis
Mikaël Chelli, Jules Descamps, Vincent Lavoué, Christophe Trojani, Michel Azar, Marcel Deckert et al. · Journal of Medical Internet Research · 2024
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.2196/53164
Methodology & findings
Study design
Comparative empirical evaluation using sequential design.
Sample
N = 471, 10 groups
Primary method
Chi-square tests to compare performance metrics and categorical variables (authors' nationalities, open-access status) between LLMs. Significance threshold P<.05. Metrics calculated: recall (TP/(TP+FN)), precision (TP/(TP+FP)), F1-score (2×precision×recall/(precision+recall)). Hallucination rate calculated as proportion of LLM-generated references with ≥2 incorrect data points among title, first author, or year. Statistical analysis performed with EasyMedStat (version 3.24).
Main result
The study found that "Precision rates for GPT-3.5, GPT-4, and Bard were 9.4% (13/139), 13.4% (16/119), and 0% (0/104) respectively (P<.001). Recall rates were 11.9% (13/109) for GPT-3.5 and 13.7% (15/109) for GPT-4, with Bard failing to retrieve any relevant papers (P<.001). Hallucination rates stood at 39.6% (55/139) for GPT-3.5, 28.6% (34/119) for GPT-4, and 91.4% (95/104) for Bard (P<.001)." The most important finding is that "using LLMs such as ChatGPT and Bard to conduct systematic reviews for a common condition such as rotator cuff disease can generate misleading or 'hallucinated' references, exceeding a 25% rate."
Reports effect sizes.
Research paradigm
Empirical-positivist (comparative measurement of LLM performance against gold-standard systematic review references)
Author conclusions
"ChatGPT and Bard exhibit the capacity to generate convincingly authentic references for systematic reviews but also yield hallucinated papers in 28.6% (34/119) to 91.3% (95/104) of cases. Among the models tested, GPT-4 displayed superior performance in generating legitimate and relevant references but, like the other models, largely failed to respect the established eligibility criteria. Given their current state, LLMs such as ChatGPT and Bard should not be used as the sole or primary means for conducting systematic reviews of literature, and it is crucial that references generated by these tools undergo rigorous validation by the authors of scientific papers."
Risk of bias
Selection bias: Study restricted to shoulder rotator cuff pathology; generalizability to other medical fields unknown; Model selection bias: Only 3 LLMs tested; other models may perform differently; Prompt engineering bias: Study tested 2 prompt versions but acknowledged alternative prompts could yield different results; Gold standard bias: Original systematic reviews serve as ground truth, but may themselves contain errors or omissions; Publication bias: Study focuses on published papers only; grey literature not included in assessment; Geographic bias: LLMs demonstrated preference for American authors (44% vs 16.5% in original reviews, P<.001); Open-access bias: LLMs preferentially retrieved open-access papers (38% vs 27.5% in originals); Temporal bias: Search restricted to 2020 publications, limiting applicability to current LLM capabilities; Selection bias: Study focused exclusively on shoulder rotator cuff pathology, limiting generalizability; Model selection bias: Only 3 LLMs tested; landscape continually evolving; Prompt bias: Authors acknowledge that choice of prompt plays crucial role; findings may not generalize to alternative query formulations; Publication bias: LLMs may have biases from training data; study found overrepresentation of American authors (44% in GPT-3.5 vs 16.5% in original reviews); Confirmation bias: Authors had clinical expertise in rotator cuff pathology, potentially influencing interpretations; Selection bias: Only shoulder rotator cuff pathology reviews were included; Model-specific bias: Only 3 LLM models tested; Prompt bias: Choice of prompt formulation may influence results; Geographical bias: LLMs over-represented American authors (44% vs 16.5% in original reviews); Open-access bias: LLMs preferentially selected open-access papers; Training data bias: LLMs trained on biased datasets may perpetuate stereotypes
Limitations
- "The scope of the study was exclusively focused on systematic reviews related to shoulder rotator cuff pathology
- Consequently, it must be recognized that the findings might not be universally applicable across diverse medical specialties or disciplines." Additionally, "The examination was also restricted to 3 LLMs, specifically GPT-3.5, GPT-4, and Bard" and "the field lacks established guidelines for leveraging LLMs to optimize accuracy
- Notwithstanding rigorous attempts to devise specific, comprehensive prompts, it remains plausible that alternative queries could generate more precise outcomes." The authors also note "Our decision not to provide the initial PubMed results list to LLMs for assessing paper eligibility was deliberate, aimed at preserving study integrity and interpretability
- While providing the list might enhance LLM accuracy, it introduces bias by guiding models toward replicating the provided set rather than autonomously identifying relevant studies."
Open questions raised
- Need for established guidelines on optimizing LLM prompts for academic tasks; further investigation needed across diverse medical fields to assess definitively whether LLMs introduce geographic and demographic biases; necessity for refining LLM training and functionality before deployment in rigorous academic purposes; need for scholarly usage guidelines integrated into LLM software itself to outline lack of liability for citation inaccuracies
- The authors identify needs for: (1) further investigation across diverse medical fields to ascertain whether LLMs introduce geographical biases definitively; (2) established guidelines for leveraging LLMs to optimize accuracy; (3) refinement of LLM training and functionality before confident use for rigorous academic purposes; (4) prominent warning statements or scholarly usage guidelines integrated into LLM software itself before deployment.
- The authors identify the need for "further investigation across diverse medical fields is warranted to ascertain whether these LLMs may introduce such biases definitively." They also note the field "lacks established guidelines for leveraging LLMs to optimize accuracy" and call for future work on optimizing prompts and developing standards for LLM use in systematic reviews.
Explore related topics
Related papers
- PRISMA Extension for Scoping Reviews (PRISMA-ScR): Checklist and ExplanationAndrea C. Tricco · 2018 · 40,391 citations
- Rayyan—a web and mobile app for systematic reviewsMourad Ouzzani · 2016 · 24,664 citations
- Cochrane Handbook for Systematic Reviews of Interventions2019 · 14,420 citations
- Guidance for conducting systematic scoping reviewsMicah D.J. Peters · 2015 · 7,472 citations
- Updated methodological guidance for the conduct of scoping reviewsMicah D.J. Peters · 2020 · 6,688 citations
- Which academic search systems are suitable for systematic reviews or meta‐analyses? Evaluating retrieval qualities of Google Scholar, PubMed, and 26 other resourcesMichael Gusenbauer · 2019 · 2,116 citations