Performance of AI Tools in Citing Retracted Literature : Content Analysis
Sebastian Labenbacher, Maximilian Niederer, Sascha Hammer, Matthias Bader, Nikolaus Schreiber, Helmar Bornemann-Cimenti · Journal of Medical Internet Research · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.2196/88766
Methodology & findings
Study design
Pragmatic cross-sectional evaluation using predefined questions (n=5) applied to 15 retracted articles (10 most-cited, 5 most-recently retracted from Retraction Watch).
Sample
N = 15, 4 groups
Primary method
Descriptive statistics including frequencies, percentages, and proportions. Cohen's kappa coefficient (κ) used to assess interreviewer agreement. Comparisons between general-purpose and research-focused AI tools reported as median values with ranges. Data extraction performed using Microsoft Excel 2023; analysis conducted using R (version 4.3.2). No formal hypothesis testing was conducted; between-tool comparisons treated as exploratory.
Main result
None of the evaluated models consistently achieved full accuracy. The study found that "none of the evaluated models consistently achieved full accuracy" and that "even the best-performing system, ChatGPT 5, correctly processed less than two-thirds of retracted articles (8 of 15)". Additionally, "AI tools marketed specifically for scientific use-Consensus, SciSpace, and ScienceOS-performed particularly poorly, each failing to produce a single fully correct set of answers."
Reports effect sizes.
Research paradigm
Positivist/empiricist - quantitative measurement of AI system performance
Risk of bias
Selection bias: Sample limited to 15 retracted articles (10 most-cited + 5 most recent); Temporal bias: Snapshot of freely accessible AI tools at mid-2025; excludes subscription models; Domain bias: 86.7% of articles were medical/biomedical, limiting generalizability; Keyword generation bias: Keywords were generated using ChatGPT 4, which may have advantaged ChatGPT in subsequent evaluations; Measurement bias: Binary scoring may mask nuanced performance differences; Selection bias: Only 15 retracted articles tested (small sample from retraction landscape); Temporal bias: Snapshot of performance at mid-2025; rapid AI model evolution limits generalizability; Domain bias: 86.7% biomedical articles may not represent performance across other scientific fields; Keyword generation bias: Authors acknowledge 'ChatGPT may have had a benefit as the keywords used for our search were generated with ChatGPT 4'; Prompt construction bias: Standardized prompt format may not reflect real-world user behavior; Tool version bias: Only free-access versions evaluated; subscription models excluded; Selection bias in retracted article selection (10 most-cited articles plus 5 most recently retracted); Potential ChatGPT bias in keyword generation (ChatGPT 4 generated keywords for articles, then ChatGPT was also evaluated); Limited sample size (n=15 retracted articles) underpowered to detect small-to-moderate differences; Temporal snapshot bias (performance evaluated mid-2025, may not reflect future iterations); Predominance of biomedical publications (86.7%) limits generalizability; Researcher bias in rating equivocal responses as incorrect (conservative scoring approach)
Limitations
- The study was "confined to freely accessible versions of each AI tool, excluding subscription models that may have more advanced retrieval capabilities and was taken at a specific point in time." Additionally, "while the test set was broad, it remains a small subset of the retraction landscape and may not generalize to all fields, especially since the majority of articles (13/15, 86.7%) were medical articles." The authors also note "the binary scoring approach (correct vs incorrect) captures factual accuracy but does not account for nuances such as ambiguous phrasing." Furthermore, they acknowledge that "our way of questioning needed a uniform, comparable way
- This was achieved by combining the keywords in a single prompt string, which could differ from a prompt by a common user searching for a specific topic."
Open questions raised
- Need for expanded sample sizes of retracted articles beyond the 15 tested
- Evaluation of subscription-based AI models with potentially more advanced retrieval capabilities
- Exploration of automated cross-validation with bibliometric databases
- Testing across non-biomedical scientific fields with differing citation and retraction practices
- Systematic benchmarking and transparent reporting of factual reliability for GenAI systems
- Development of retraction-aware pipelines with dynamic metadata synchronization
Explore related topics
Related papers
- Estimating the reproducibility of psychological scienceAlexander A. Aarts · 2015 · 8,669 citations
- Autonomous chemical research with large language modelsDaniil A. Boiko · 2023 · 809 citations
- Academic Integrity considerations of AI Large Language Models in the post-pandemic era: ChatGPT and beyondMike Perkins · 2023 · 668 citations
- Comparing scientific abstracts generated by ChatGPT to real abstracts with detectors and blinded human reviewersCatherine A. Gao · 2023 · 657 citations
- Teacher support and student motivation to learn with Artificial Intelligence (AI) based chatbotThomas K. F. Chiu · 2023 · 617 citations
- ChatGPT and the rise of generative AI: Threat to academic integrity?Damian Eke · 2023 · 476 citations