12,637 papers · updated 18 Sept 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Can generative AI reliably synthesise literature? exploring hallucination issues in ChatGPT

Amr Adel, Noor Haitham Saleem · AI & Society · 2025

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
0/4
Quality (LMQS)
E
Evidence
30
Citations
12.85
FWCI
Top 10%
Impact

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1007/s00146-025-02406-7

Methodology & findings

Study design

Systematic literature review following PRISMA guidelines with a case study component.

Sample

N = 124, 3 groups

Primary method

Qualitative synthesis and thematic analysis of included studies. Quantitative metrics extracted include sensitivity, specificity, precision, hallucination rates, and time efficiency measures. Studies comparing AI to human reviewers were categorized (8 reported comparable performance, 3 found AI underperformed, 2 noted superior AI performance). An eight-point scoring rubric was applied independently by AI (via Elicit) and human reviewers, with consensus resolution of discrepancies.

Main result

The study found that "ChatGPT exhibited strong performance in structured tasks, such as title and abstract screening, especially when using GPT-4" with "reported sensitivity ranged from 80.6% to 96.5%," yet "hallucination remained a critical concern" with rates "between 28 and 91%, highlighting the model's tendency to fabricate plausible-sounding references or facts." Additionally, "efficiency metrics quantified across studies demonstrate substantial time savings, ranging between 40 to 90%" compared to traditional methods, though "performance declined in more interpretative contexts, with precision dropping as low as 4.6%."

Reports effect sizes and confidence intervals.

Research paradigm

Critical realism with pragmatist orientation; examines technology reliability through empirical assessment while acknowledging sociotechnical contexts

Author conclusions

"Results indicate that generative AI models can substantially improve efficiency, with time savings of up to 40% across various review tasks" and "Accuracy was particularly strong in structured fields such as medical research, where title and abstract screening sensitivities ranged from 80.6% to 96.2%. However, performance declined in more interpretative contexts, with precision dropping as low as 4.6% and hallucination rates reaching 91%." The authors conclude: "These findings underscore the value of AI-human collaboration, where human oversight is essential to ensure methodological rigour and prevent misinformation in academic synthesis." They emphasize that "Most reviewed studies recommend hybrid human-AI models, underscoring collaborative frameworks as optimal for ensuring research accuracy and integrity."

Risk of bias

Selection bias: Limited availability of peer-reviewed studies in nascent field; Publication bias: Paywalled journal access restrictions; Search engine bias: Limited diversity in initial database searches; Reporting bias: Lack of standardized reporting across included studies; Selection bias in study inclusion: Manual screening may introduce subjective judgment despite rigorous criteria; Selection bias: manual screening applied to address initial search limitations; Non-peer-reviewed sources risk misinformation; Task-specific performance variability not standardized across studies; Hallucination under-reporting in literature; Publication bias: reliance on published peer-reviewed studies; preprints and grey literature supplemented but may not capture all relevant work; Screening bias: discrepancies between AI-assisted and manual screening not fully controlled; many sources discussed AI broadly without focusing on literature synthesis; Model version bias: Performance differences across GPT-3.5, ChatGPT 3.5, and GPT-4 versions; some sources unavailable; manual verification performed only after selection; Selection bias in choosing high-quality studies (only 40 of 124 reviewed in depth)

Limitations

  • "Despite a rigorous search strategy, challenges emerged due to the limited availability of peer-reviewed studies specifically addressing AI-assisted literature reviews, as the field remains nascent." Additional limitations included: "Paywalled journals restricted access to some high-quality works" and "non-peer-reviewed articles risked misinformation." The authors note "the lack of standardised reporting made comparisons across studies difficult," and crucially, "Even when outputs were accurate, their rationale often remained opaque, complicating verification and eroding trust in high-stakes settings." The study also acknowledges that hallucination may be "under-reported in others" beyond the three studies explicitly identified.

Open questions raised

  • Enhancing ChatGPT's ability to handle complex logical structures and nuanced academic debates through domain-specific fine-tuning
  • Establishing standardised evaluation frameworks to objectively benchmark AI-assisted literature review tasks
  • Cross-disciplinary comparative studies to ascertain AI's generalisability across diverse academic fields
  • Exploring advanced hybrid human-AI workflows that leverage AI efficiency while maintaining scholarly oversight and ethical transparency
  • Understanding performance variability across datasets and disciplinary contexts
  • Investigating ChatGPT's performance in interpretative contexts beyond medical research
Extracted from: pdfAgreement 57%

Explore related topics

Related papers