Does ChatGPT Ignore Article Retractions and Other Reliability Concerns?
Mike Thelwall, Marianna Lehtisaari, Irini Katsirea, Kim Holmberg · Learned Publishing · 2025
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1002/leap.2018
Methodology & findings
Study design
Mixed-methods empirical study with two research questions: (1) RQ1 used a quality evaluation task where 217 high-profile retracted or potentially concerning articles (selected from 56,705 retracted/concerning articles based on Altmetric scores for media mentions) had their titles and abstracts submitted 30 times each to ChatGPT 4o-mini via API, requesting quality scores using REF 2021 guidelines.
Sample
N = 217, 13 groups
Primary method
Descriptive statistics: mean ChatGPT quality scores (averaged across 30 iterations per article); Correlation analysis: Spearman's rank correlation with bootstrap confidence intervals (95% CI); Inter-rater reliability: Cohen's Kappa for coding agreement between two raters; Manual content analysis with categorical coding scheme (9 categories for truth-value assessment); Confusion matrix analysis for coder disagreement; Frequency analysis and percentage calculations for RQ2 claim categorization
Main result
The study found that ChatGPT 4o-mini did not demonstrate awareness of article retractions when evaluating research quality. Specifically, "None of the 217 × 30 ChatGPT reports on the quality of the 217 articles mentioned retractions, corrections, or ethical problems." Additionally, for claims extracted from retracted articles, "almost two-thirds of the time, ChatGPT reported that they were likely to be true (43.8%), partially true (9.5%), or consistent with research (9.5%)." The study concludes that "ChatGPT does not have awareness of retractions or other indicators of problematic content for high-profile academic research."
Reports effect sizes and confidence intervals.
Research paradigm
Empirical-positivist
Author conclusions
"The results overall suggest that ChatGPT does not have awareness of retractions or other indicators of problematic content for high-profile academic research, although it is sensitive to some particularly important health issues associated with retractions. This lack of awareness seems to be an unfortunate limitation that should be addressed in the future. While people visiting retracted article pages would now typically see clear retraction notices, if they rely instead on LLMs for knowledge or knowledge summaries, then they can be misled." The authors emphasize that this represents "only one issue of false information in LLM output, but as retracted articles seem now to be clearly flagged, it seems unfortunate that this issue does not seem to have been addressed yet by ChatGPT."
Risk of bias
Selection bias: High-profile articles (based on Altmetric scores) may not be representative of all retracted articles; Temporal bias: ChatGPT knowledge cutoff of October 1, 2023 may exclude recent retractions; only 8 articles (4%) had first serious concerns raised after cutoff; Source bias: RetractionWatch may contain unknown biases in which articles are flagged; Coder bias: Initial inter-rater reliability was moderate (Cohen's Kappa = 0.598), requiring first author to reconcile disagreements; Confounding: More well-known articles might score higher due to genuine research quality independent of retraction status; Selection bias: Study limited to high-profile retractions (Altmetric-ranked), excluding lower-profile retractions; Temporal bias: ChatGPT training cutoff of October 1, 2023 may exclude recent retractions (though only 4% of analyzed articles affected); Operationalization bias: 'High-profile' defined by online mentions (Altmetric scores), not necessarily by scientific importance or retraction severity; Coding bias: Inter-rater reliability moderate (Kappa=0.598), with disagreement resolution by first author only; Source bias: Reliance on RetractionWatch as primary data source acknowledged as potentially biased; Duplicate article bias: Acknowledged imperfect duplicate detection may have left duplicates in 56,705-article initial dataset; Incomplete abstract data: 31 of 250 top-ranked articles excluded due to missing abstracts; Confounding: More well-known articles might score higher due to importance rather than retraction awareness; Selection bias: Articles selected based on Altmetric scores may over-represent well-publicized cases and not represent all retractions; Knowledge cutoff bias: ChatGPT training cutoff of October 1, 2023, means retractions after this date may not be detected; Source bias: RetractionWatch database may contain unknown biases in which articles it flags as concerning; Measurement bias: Altmetric scores may not accurately represent actual prominence or importance of articles; Coding bias: Although Cohen's Kappa = 0.598 achieved moderate agreement, disagreements were resolved by first author only
Limitations
- "This study is limited to a single LLM (ChatGPT 4o-mini), a single set of high-profile articles, a single period (September 2024), and the UK REF definition of research quality
- It is possible that other LLMs will work differently, and it seems likely that LLM processing and reporting of scientific information will evolve over time." Additionally, "Since the ChatGPT 4o-mini version used had knowledge ending on October 1, 2023, retractions and concerns expressed after this date might be unknown to it." The authors also note that "The main source of retracted or potentially problematic articles, RetractionWatch, is also a limitation since it may contain unknown biases."
Open questions raised
- Whether other LLMs (Gemini, DeepSeek) handle retractions differently than ChatGPT
- How LLM processing and reporting of scientific information evolves over time as newer versions integrate live web sources
- Whether different prompt strategies could reveal ChatGPT's knowledge of retractions
- Comprehensive verification of complex retracted claims across all 61 statements
- Whether newer LLM versions (post-June 2025) with live web integration address retraction processing
- Whether other LLMs (Gemini, DeepSeek, etc.) handle retracted articles differently than ChatGPT
Explore related topics
Related papers
- Estimating the reproducibility of psychological scienceAlexander A. Aarts · 2015 · 8,669 citations
- Autonomous chemical research with large language modelsDaniil A. Boiko · 2023 · 809 citations
- Academic Integrity considerations of AI Large Language Models in the post-pandemic era: ChatGPT and beyondMike Perkins · 2023 · 668 citations
- Comparing scientific abstracts generated by ChatGPT to real abstracts with detectors and blinded human reviewersCatherine A. Gao · 2023 · 657 citations
- Teacher support and student motivation to learn with Artificial Intelligence (AI) based chatbotThomas K. F. Chiu · 2023 · 617 citations
- ChatGPT and the rise of generative AI: Threat to academic integrity?Damian Eke · 2023 · 476 citations