12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

On the utility of ChatGPT in conducting a literature review on deep learning for dopamine transporter SPECT with [¹²³I]ioflupane

Henning Boecker, Ralph Buchert, Thomas Buddenkotte, Ivayla Apostolova, Roland Opfer, Susanne Klutmann et al. · Nuklearmedizin - NuclearMedicine · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
0/4
Quality (LMQS)
E
Evidence
0
Citations
0.00
FWCI

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1055/a-2890-8325

Methodology & findings

Study design

Multi-step workflow evaluation with manual fact-checking: (i) literature search using ChatGPT, (ii) generation of structured summaries with 24 predefined fields per publication using ChatGPT with iteratively designed prompts, and (iii) drafting a review paper using ChatGPT.

Sample

N = 67, 6 groups

Primary method

Descriptive statistics: proportions (percentage of summaries requiring corrections: 40.3% with corrections, 59.7% without corrections). No inferential statistical tests reported.

Main result

The study found that "ChatGPT cited 13 papers, whereas the manual search identified 70 relevant publications, 67 of which were included." Additionally, "Corrections to ChatGPT-generated structured summaries were required in 27 cases (40.3%), affecting one or two of the 24 predefined fields, while no changes were necessary in 40 publications (59.7%)." The authors note that "All numerical information (dataset sizes, train-test splits, performance metrics) was correct." and "The review draft (~950 words) generated by ChatGPT was content-wise meaningful and accurate, but contained referencing errors, including incorrect citations, missing references, and citations of non-existent publications."

Reports effect sizes.

Research paradigm

Empiricist/pragmatist (evaluation of tool utility)

Author conclusions

The authors conclude that "ChatGPT is a highly effective tool for drafting review manuscripts in nuclear medicine imaging, but its limitations in literature retrieval and referencing require careful expert supervision."

Risk of bias

ChatGPT's known limitations in exhaustive literature retrieval; Potential for hallucination in citations and references; Dependence on manual expert fact-checking and correction; Single expert conducting manual validation; Single expert performer for manual literature search (no inter-rater reliability checks); Potential selection bias in which papers ChatGPT could access or retrieve; No blinding during fact-checking process; ChatGPT version specification (GPT-5.2) is unusually futuristic for a 2026 publication, suggesting possible metadata error; Single domain focus (DAT-SPECT and deep learning) limits generalizability claims; Selection bias: Single expert performed manual literature search, potentially missing relevant publications or introducing personal bias; Tool-specific bias: ChatGPT version (GPT-5.2) may have inherent training biases regarding literature retrieval; Verification bias: Manual fact-checking by study team may not catch all errors; Prompt design bias: Iterative prompt refinement with ChatGPT support may have optimized prompts beyond typical user capability

Limitations

  • The study indicates that "ChatGPT is a highly effective tool for drafting review manuscripts in nuclear medicine imaging, but its limitations in literature retrieval and referencing require careful expert supervision." The abstract notes that ChatGPT demonstrated significant gaps in literature retrieval (retrieving only 13 of 70 relevant papers) and produced referencing errors including "incorrect citations, missing references, and citations of non-existent publications."
Data: not_statedCode: not_statedExtracted from: pdfAgreement 70%

Explore related topics

Related papers