Can generative AI reliably synthesise literature? exploring hallucination issues in ChatGPT
Amr Adel, Noor Haitham Saleem · AI & Society · 2025
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1007/s00146-025-02406-7
Methodology & findings
Study design
Systematic literature review following PRISMA guidelines with a case study component.
Sample
N = 124, 3 groups
Primary method
Qualitative synthesis and thematic analysis of included studies. Quantitative metrics extracted include sensitivity, specificity, precision, hallucination rates, and time efficiency measures. Studies comparing AI to human reviewers were categorized (8 reported comparable performance, 3 found AI underperformed, 2 noted superior AI performance). An eight-point scoring rubric was applied independently by AI (via Elicit) and human reviewers, with consensus resolution of discrepancies.
Main result
The study found that "ChatGPT exhibited strong performance in structured tasks, such as title and abstract screening, especially when using GPT-4" with "reported sensitivity ranged from 80.6% to 96.5%," yet "hallucination remained a critical concern" with rates "between 28 and 91%, highlighting the model's tendency to fabricate plausible-sounding references or facts." Additionally, "efficiency metrics quantified across studies demonstrate substantial time savings, ranging between 40 to 90%" compared to traditional methods, though "performance declined in more interpretative contexts, with precision dropping as low as 4.6%."
Reports effect sizes and confidence intervals.
Research paradigm
Critical realism with pragmatist orientation; examines technology reliability through empirical assessment while acknowledging sociotechnical contexts
Author conclusions
"Results indicate that generative AI models can substantially improve efficiency, with time savings of up to 40% across various review tasks" and "Accuracy was particularly strong in structured fields such as medical research, where title and abstract screening sensitivities ranged from 80.6% to 96.2%. However, performance declined in more interpretative contexts, with precision dropping as low as 4.6% and hallucination rates reaching 91%." The authors conclude: "These findings underscore the value of AI-human collaboration, where human oversight is essential to ensure methodological rigour and prevent misinformation in academic synthesis." They emphasize that "Most reviewed studies recommend hybrid human-AI models, underscoring collaborative frameworks as optimal for ensuring research accuracy and integrity."
Risk of bias
Selection bias: Limited availability of peer-reviewed studies in nascent field; Publication bias: Paywalled journal access restrictions; Search engine bias: Limited diversity in initial database searches; Reporting bias: Lack of standardized reporting across included studies; Selection bias in study inclusion: Manual screening may introduce subjective judgment despite rigorous criteria; Selection bias: manual screening applied to address initial search limitations; Non-peer-reviewed sources risk misinformation; Task-specific performance variability not standardized across studies; Hallucination under-reporting in literature; Publication bias: reliance on published peer-reviewed studies; preprints and grey literature supplemented but may not capture all relevant work; Screening bias: discrepancies between AI-assisted and manual screening not fully controlled; many sources discussed AI broadly without focusing on literature synthesis; Model version bias: Performance differences across GPT-3.5, ChatGPT 3.5, and GPT-4 versions; some sources unavailable; manual verification performed only after selection; Selection bias in choosing high-quality studies (only 40 of 124 reviewed in depth)
Limitations
- "Despite a rigorous search strategy, challenges emerged due to the limited availability of peer-reviewed studies specifically addressing AI-assisted literature reviews, as the field remains nascent." Additional limitations included: "Paywalled journals restricted access to some high-quality works" and "non-peer-reviewed articles risked misinformation." The authors note "the lack of standardised reporting made comparisons across studies difficult," and crucially, "Even when outputs were accurate, their rationale often remained opaque, complicating verification and eroding trust in high-stakes settings." The study also acknowledges that hallucination may be "under-reported in others" beyond the three studies explicitly identified.
Open questions raised
- Enhancing ChatGPT's ability to handle complex logical structures and nuanced academic debates through domain-specific fine-tuning
- Establishing standardised evaluation frameworks to objectively benchmark AI-assisted literature review tasks
- Cross-disciplinary comparative studies to ascertain AI's generalisability across diverse academic fields
- Exploring advanced hybrid human-AI workflows that leverage AI efficiency while maintaining scholarly oversight and ethical transparency
- Understanding performance variability across datasets and disciplinary contexts
- Investigating ChatGPT's performance in interpretative contexts beyond medical research
Explore related topics
Related papers
- Preferred reporting items for systematic reviews and meta-analyses: the PRISMA statementDavid Moher · 2009 · 83,271 citations
- PRISMA Extension for Scoping Reviews (PRISMA-ScR): Checklist and ExplanationAndrea C. Tricco · 2018 · 40,391 citations
- Rayyan—a web and mobile app for systematic reviewsMourad Ouzzani · 2016 · 24,664 citations
- Cochrane Handbook for Systematic Reviews of Interventions2019 · 14,420 citations
- PRISMA 2020 explanation and elaboration: updated guidance and exemplars for reporting systematic reviewsMatthew J. Page · 2021 · 10,956 citations
- Guidance for conducting systematic scoping reviewsMicah D.J. Peters · 2015 · 7,472 citations