12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Can generative AI reliably synthesise literature? exploring hallucination issues in ChatGPT

Amr Adel, Noor Haitham Saleem · AI & Society · 2025

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
0/4
Quality (LMQS)
E
Evidence
30
Citations
12.85
FWCI
Top 10%
Impact

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1007/s00146-025-02406-7

Methodology & findings

Study design

Systematic literature review following PRISMA guidelines with a case study component.

Sample

N = 124, 9 groups

Primary method

Qualitative synthesis and thematic analysis of included studies. Quantitative metrics extracted include sensitivity, specificity, precision, hallucination rates, and time efficiency measures. Studies comparing AI to human reviewers were categorized (8 reported comparable performance, 3 found AI underperformed, 2 noted superior AI performance). An eight-point scoring rubric was applied independently by AI (via Elicit) and human reviewers, with consensus resolution of discrepancies.

Main result

The study found that "ChatGPT exhibited strong performance in structured tasks, such as title and abstract screening, especially when using GPT-4" with "reported sensitivity ranged from 80.6% to 96.5%," yet "hallucination remained a critical concern" with rates "between 28 and 91%, highlighting the model's tendency to fabricate plausible-sounding references or facts." Additionally, "efficiency metrics quantified across studies demonstrate substantial time savings, ranging between 40 to 90%" compared to traditional methods, though "performance declined in more interpretative contexts, with precision dropping as low as 4.6%."

Reports effect sizes and confidence intervals.

Research paradigm

Critical realism with pragmatist orientation; examines technology reliability through empirical assessment while acknowledging sociotechnical contexts

Author conclusions

"Results indicate that generative AI models can substantially improve efficiency, with time savings of up to 40% across various review tasks" and "Accuracy was particularly strong in structured fields such as medical research, where title and abstract screening sensitivities ranged from 80.6% to 96.2%. However, performance declined in more interpretative contexts, with precision dropping as low as 4.6% and hallucination rates reaching 91%." The authors conclude: "These findings underscore the value of AI-human collaboration, where human oversight is essential to ensure methodological rigour and prevent misinformation in academic synthesis." They emphasize that "Most reviewed studies recommend hybrid human-AI models, underscoring collaborative frameworks as optimal for ensuring research accuracy and integrity."

Risk of bias

Selection bias: Limited availability of peer-reviewed studies in nascent field; Publication bias: Paywalled journal access restrictions; Search engine bias: Limited diversity in initial database searches; Reporting bias: Lack of standardized reporting across included studies; Selection bias in study inclusion: Manual screening may introduce subjective judgment despite rigorous criteria; Publication bias: limited peer-reviewed literature on nascent field; Paywalled article access restriction; Search engine bias limiting diversity in initial retrieval; Selection bias: manual screening applied to address initial search limitations; Non-peer-reviewed sources risk misinformation; Task-specific performance variability not standardized across studies; Hallucination under-reporting in literature; Database bias: search engine limitations and paywalled journal access restrictions; Publication bias: reliance on published peer-reviewed studies; preprints and grey literature supplemented but may not capture all relevant work; Study heterogeneity: significant variability in reporting standards across included studies limited direct comparisons; Screening bias: discrepancies between AI-assisted and manual screening not fully controlled; Selection bias: Limited availability of peer-reviewed studies on AI-assisted literature reviews due to field nascency; Search bias: Paywalled journals restricting access; search engine bias limiting diversity; Reporting bias: Variability in standardised reporting across included studies; hallucination rates likely under-reported; Publication bias: Non-peer-reviewed articles risked misinformation; many sources discussed AI broadly without focusing on literature synthesis; Model version bias: Performance differences across GPT-3.5, ChatGPT 3.5, and GPT-4 versions; Selection bias: Limited availability of peer-reviewed studies in nascent field; many sources discussed AI broadly without focusing on literature synthesis; Database bias: Search engine bias limited diversity in study selection; Access bias: Paywalled journals restricted access to high-quality works; some sources unavailable; Reporting bias: Hallucination rates likely under-reported across studies; lack of standardized reporting metrics; Publication bias: Reliance on peer-reviewed sources may exclude relevant grey literature findings; Information bias: Non-peer-reviewed articles at risk of misinformation; manual verification performed only after selection; Screening bias: Potential discrepancies between AI-assisted and manual screening in Elicit recommendation phase; Model variation bias: Performance variability across GPT versions (GPT-3.5 vs GPT-4) and models (ChatGPT, ChatSonic, Bard) confounds interpretation; Search engine bias limiting diversity in initial database searches; Publication bias (paywalled journals restricting access); Potential misinformation from non-peer-reviewed articles; Lack of standardized reporting across included studies; Selection bias in choosing high-quality studies (only 40 of 124 reviewed in depth)

Limitations

  • "Despite a rigorous search strategy, challenges emerged due to the limited availability of peer-reviewed studies specifically addressing AI-assisted literature reviews, as the field remains nascent." Additional limitations included: "Paywalled journals restricted access to some high-quality works" and "non-peer-reviewed articles risked misinformation." The authors note "the lack of standardised reporting made comparisons across studies difficult," and crucially, "Even when outputs were accurate, their rationale often remained opaque, complicating verification and eroding trust in high-stakes settings." The study also acknowledges that hallucination may be "under-reported in others" beyond the three studies explicitly identified.

Open questions raised

  • Enhancing ChatGPT's ability to handle complex logical structures and nuanced academic debates through domain-specific fine-tuning
  • Establishing standardised evaluation frameworks to objectively benchmark AI-assisted literature review tasks
  • Cross-disciplinary comparative studies to ascertain AI's generalisability across diverse academic fields
  • Exploring advanced hybrid human-AI workflows that leverage AI efficiency while maintaining scholarly oversight and ethical transparency
  • Understanding performance variability across datasets and disciplinary contexts
  • Establishing standardized evaluation frameworks to objectively benchmark AI-assisted literature review tasks
Extracted from: pdfAgreement 57%

Explore related topics

Related papers