12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Does ChatGPT Ignore Article Retractions and Other Reliability Concerns?

Mike Thelwall, Marianna Lehtisaari, Irini Katsirea, Kim Holmberg · Learned Publishing · 2025

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
1/4
Quality (LMQS)
E
Evidence
11
Citations
4.67
FWCI
Top 10%
Impact

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1002/leap.2018

Methodology & findings

Study design

Mixed-methods empirical study with two research questions: (1) RQ1 used a quality evaluation task where 217 high-profile retracted or potentially concerning articles (selected from 56,705 retracted/concerning articles based on Altmetric scores for media mentions) had their titles and abstracts submitted 30 times each to ChatGPT 4o-mini via API, requesting quality scores using REF 2021 guidelines.

Sample

N = 217, 13 groups

Primary method

Descriptive statistics: mean ChatGPT quality scores (averaged across 30 iterations per article); Correlation analysis: Spearman's rank correlation with bootstrap confidence intervals (95% CI); Inter-rater reliability: Cohen's Kappa for coding agreement between two raters; Manual content analysis with categorical coding scheme (9 categories for truth-value assessment); Confusion matrix analysis for coder disagreement; Frequency analysis and percentage calculations for RQ2 claim categorization

Main result

The study found that ChatGPT 4o-mini did not demonstrate awareness of article retractions when evaluating research quality. Specifically, "None of the 217 × 30 ChatGPT reports on the quality of the 217 articles mentioned retractions, corrections, or ethical problems." Additionally, for claims extracted from retracted articles, "almost two-thirds of the time, ChatGPT reported that they were likely to be true (43.8%), partially true (9.5%), or consistent with research (9.5%)." The study concludes that "ChatGPT does not have awareness of retractions or other indicators of problematic content for high-profile academic research."

Reports effect sizes and confidence intervals.

Research paradigm

Empirical-positivist

Author conclusions

"The results overall suggest that ChatGPT does not have awareness of retractions or other indicators of problematic content for high-profile academic research, although it is sensitive to some particularly important health issues associated with retractions. This lack of awareness seems to be an unfortunate limitation that should be addressed in the future. While people visiting retracted article pages would now typically see clear retraction notices, if they rely instead on LLMs for knowledge or knowledge summaries, then they can be misled." The authors emphasize that this represents "only one issue of false information in LLM output, but as retracted articles seem now to be clearly flagged, it seems unfortunate that this issue does not seem to have been addressed yet by ChatGPT."

Risk of bias

Selection bias: High-profile articles (based on Altmetric scores) may not be representative of all retracted articles; Temporal bias: ChatGPT knowledge cutoff of October 1, 2023 may exclude recent retractions; only 8 articles (4%) had first serious concerns raised after cutoff; Source bias: RetractionWatch may contain unknown biases in which articles are flagged; Coder bias: Initial inter-rater reliability was moderate (Cohen's Kappa = 0.598), requiring first author to reconcile disagreements; Confounding: More well-known articles might score higher due to genuine research quality independent of retraction status; Selection bias: Study limited to high-profile retractions (Altmetric-ranked), excluding lower-profile retractions; Temporal bias: ChatGPT training cutoff of October 1, 2023 may exclude recent retractions (though only 4% of analyzed articles affected); Operationalization bias: 'High-profile' defined by online mentions (Altmetric scores), not necessarily by scientific importance or retraction severity; Coding bias: Inter-rater reliability moderate (Kappa=0.598), with disagreement resolution by first author only; Source bias: Reliance on RetractionWatch as primary data source acknowledged as potentially biased; Duplicate article bias: Acknowledged imperfect duplicate detection may have left duplicates in 56,705-article initial dataset; Incomplete abstract data: 31 of 250 top-ranked articles excluded due to missing abstracts; Confounding: More well-known articles might score higher due to importance rather than retraction awareness; Selection bias: Articles selected based on Altmetric scores may over-represent well-publicized cases and not represent all retractions; Knowledge cutoff bias: ChatGPT training cutoff of October 1, 2023, means retractions after this date may not be detected; Source bias: RetractionWatch database may contain unknown biases in which articles it flags as concerning; Measurement bias: Altmetric scores may not accurately represent actual prominence or importance of articles; Coding bias: Although Cohen's Kappa = 0.598 achieved moderate agreement, disagreements were resolved by first author only

Limitations

  • "This study is limited to a single LLM (ChatGPT 4o-mini), a single set of high-profile articles, a single period (September 2024), and the UK REF definition of research quality
  • It is possible that other LLMs will work differently, and it seems likely that LLM processing and reporting of scientific information will evolve over time." Additionally, "Since the ChatGPT 4o-mini version used had knowledge ending on October 1, 2023, retractions and concerns expressed after this date might be unknown to it." The authors also note that "The main source of retracted or potentially problematic articles, RetractionWatch, is also a limitation since it may contain unknown biases."

Open questions raised

  • Whether other LLMs (Gemini, DeepSeek) handle retractions differently than ChatGPT
  • How LLM processing and reporting of scientific information evolves over time as newer versions integrate live web sources
  • Whether different prompt strategies could reveal ChatGPT's knowledge of retractions
  • Comprehensive verification of complex retracted claims across all 61 statements
  • Whether newer LLM versions (post-June 2025) with live web integration address retraction processing
  • Whether other LLMs (Gemini, DeepSeek, etc.) handle retracted articles differently than ChatGPT
Data: Retracted and concerning articles data from CrossRef RetractionWatch Database: https://www.crossref.org/blog/news-cross-ref-and-retraction-watch/; Altmetric data: Obtained via Altmetric.com API (DOIs of 56,705 articles submitted); Retracted and potentially concerning articles dataset derived from RetractionWatch Database accessed via CrossRef (https://www.crossref.org/blog/news-cross-ref-and-retraction-watch/); Scopus retracted articles dataset (downloaded via Scopus API, DOI-matched records); Altmetric scores for 56,705 articles retrieved via Altmetric.com API; Retracted articles dataset from CrossRef (https://www.crossref.org/blog/news-cross-ref-and-retraction-watch/); Altmetric data accessed via Altmetric.com API; Scopus retracted articles dataExtracted from: pdfAgreement 57%

Explore related topics

Related papers