12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Do Large Language Models know Which Published Articles have been Retracted?

Mike Thelwall · arXiv (Cornell University) · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
1/4
Quality (LMQS)
E
Evidence
0
Citations
0.00
FWCI

Methodology & findings

Study design

Benchmark study combining qualitative and quantitative analysis.

Sample

N = 6740.5, 8 groups

Primary method

Chi-square test of independence (2×2×2 design) to assess whether different LLMs gave similar results. Simple descriptive analysis (frequencies, percentages) for overall performance metrics. Manual ad-hoc checking of error claims on subset of 30 articles.

Main result

The study found that "All three LLMs described the high-profile retracted articles as not retracted at least 82% of the time", with GPT OSS 120B identifying only 18% of retracted articles, Gemma 3 27B identifying 16%, and DeepSeek R1 70B identifying 12%. Additionally, "for the full text input, there were 63-8=55 cases of an LLM (GPT OSS 120B: 40; Gemma 3 27B: 15) stating that a non-retracted article in the benchmark dataset had been retracted, and 48,977 of the LLM stating that an article had not been retracted (an 0.11% error rate)."

Reports effect sizes.

Research paradigm

Empirical/Positivist - quantitative assessment of LLM capabilities through structured benchmarking

Author conclusions

"The results suggest that three major contemporary LLMs are rarely able to report that articles are retracted, even if they are high profile. This applies only to the open weights locally run versions since the web interfaces can run web searches to check directly. This gives additional evidence that LLM deductions about academic research must be interpreted cautiously, especially when they are used in offline mode. On the positive side, they rarely claim that non-retraced articles have been retracted, and when they do it seems safe to ignore the claims as they are unlikely to be due to errors identified with the paper."

Risk of bias

Selection bias: High-profile articles chosen to maximize LLM awareness, representing best-case scenario; Language bias: Prompts in English only; LLM selection bias: Purposive selection of three major LLMs rather than comprehensive sampling; Data source bias: Benchmark dataset derived from MDPI journals only (publisher with specific characteristics); Temporal bias: Training data cutoff dates vary across LLMs (June-August 2024); Selection bias: High-profile retracted articles selected specifically to give LLMs 'best possible chance' to have learned about retraction, creating a best-case scenario that may not generalize; Dataset bias: Benchmark dataset limited to MDPI journals only, not fully representative of all published research; Prompt bias: Results dependent on specific prompts used; different prompts may yield different results; Model selection bias: Purposive (non-random) selection of only three LLMs from tens of thousands available; Selection bias: High-profile articles chosen as 'best-case scenario' for LLM knowledge, which may not represent typical retracted articles; Sample composition bias: Articles from MDPI journals and biomedical-heavy sources may not represent all academic fields equally; Measurement bias: Reliance on structured prompt format may affect how LLMs respond; different prompting strategies could yield different results; Detection bias: LLMs may have difficulty detecting retractions in offline mode, creating artificial performance deficit compared to online-capable versions

Limitations

  • "The results are limited by the choice of LLMs and the language of the prompts
  • Different results may have been obtained from other LLMs or prompts
  • The results do not explain why LLMs rarely know that high profile articles have been retracted." Additionally, the study notes that "This applies only to the open weights locally run versions since the web interfaces can run web searches to check directly."

Open questions raised

  • Why LLMs rarely know that high-profile articles have been retracted despite having ingested this information multiple times
  • Whether alternative prompting strategies or agent-based approaches might elicit retraction information more effectively
  • Performance of web-enabled LLM interfaces that can perform online checks
  • Assessment across a broader range of LLMs beyond the three tested
  • The study does not explain why LLMs rarely know that high-profile articles have been retracted
  • Unclear whether retraction information is usually forgotten or stored in a way that does not allow connection to articles through titles and abstracts
Data: 161 high-profile retracted articles from Retraction Watch (recycled from Thelwall et al., 2025); 5,000+ articles from MDPI journals across eight fields (Frontiers and MDPI XML collections); Three retracted articles accidentally included in MDPI benchmark dataset; MDPI journal article collection (XML format, downloaded December 2025 via FTP access from MDPI); High-profile retracted articles dataset (161 articles retracted before October 2023, sourced from previous study based on Retraction Watch list); High-profile retracted articles dataset (217 articles after filtering, 161 retracted before October 2023) recycled from Thelwall et al. (2025), sourced from Retraction Watch and Altmetric.com; MDPI journal articles dataset: 8 journals with up to 5,000 articles per journal (total variable depending on journal size; full dataset downloaded via FTP from MDPI in December 2025)Code: Python program for XML-to-plain-text conversion (similar to publicly available code mentioned at Böschen, 2021); Python program for converting XML to plain text (similar to publicly available program referenced as Böschen, 2021); Python program for XML-to-plaintext conversion (similar to publicly available code referenced as Böschen, 2021)Extracted from: pdfAgreement 53%

Explore related topics

Related papers