12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Performance of AI Tools in Citing Retracted Literature : Content Analysis

Sebastian Labenbacher, Maximilian Niederer, Sascha Hammer, Matthias Bader, Nikolaus Schreiber, Helmar Bornemann-Cimenti · Journal of Medical Internet Research · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
1/4
Quality (LMQS)
E
Evidence
0
Citations
0.00
FWCI

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.2196/88766

Methodology & findings

Study design

Pragmatic cross-sectional evaluation using predefined questions (n=5) applied to 15 retracted articles (10 most-cited, 5 most-recently retracted from Retraction Watch).

Sample

N = 15, 4 groups

Primary method

Descriptive statistics including frequencies, percentages, and proportions. Cohen's kappa coefficient (κ) used to assess interreviewer agreement. Comparisons between general-purpose and research-focused AI tools reported as median values with ranges. Data extraction performed using Microsoft Excel 2023; analysis conducted using R (version 4.3.2). No formal hypothesis testing was conducted; between-tool comparisons treated as exploratory.

Main result

None of the evaluated models consistently achieved full accuracy. The study found that "none of the evaluated models consistently achieved full accuracy" and that "even the best-performing system, ChatGPT 5, correctly processed less than two-thirds of retracted articles (8 of 15)". Additionally, "AI tools marketed specifically for scientific use-Consensus, SciSpace, and ScienceOS-performed particularly poorly, each failing to produce a single fully correct set of answers."

Reports effect sizes.

Research paradigm

Positivist/empiricist - quantitative measurement of AI system performance

Risk of bias

Selection bias: Sample limited to 15 retracted articles (10 most-cited + 5 most recent); Temporal bias: Snapshot of freely accessible AI tools at mid-2025; excludes subscription models; Domain bias: 86.7% of articles were medical/biomedical, limiting generalizability; Keyword generation bias: Keywords were generated using ChatGPT 4, which may have advantaged ChatGPT in subsequent evaluations; Measurement bias: Binary scoring may mask nuanced performance differences; Selection bias: Only 15 retracted articles tested (small sample from retraction landscape); Temporal bias: Snapshot of performance at mid-2025; rapid AI model evolution limits generalizability; Domain bias: 86.7% biomedical articles may not represent performance across other scientific fields; Keyword generation bias: Authors acknowledge 'ChatGPT may have had a benefit as the keywords used for our search were generated with ChatGPT 4'; Prompt construction bias: Standardized prompt format may not reflect real-world user behavior; Tool version bias: Only free-access versions evaluated; subscription models excluded; Selection bias in retracted article selection (10 most-cited articles plus 5 most recently retracted); Potential ChatGPT bias in keyword generation (ChatGPT 4 generated keywords for articles, then ChatGPT was also evaluated); Limited sample size (n=15 retracted articles) underpowered to detect small-to-moderate differences; Temporal snapshot bias (performance evaluated mid-2025, may not reflect future iterations); Predominance of biomedical publications (86.7%) limits generalizability; Researcher bias in rating equivocal responses as incorrect (conservative scoring approach)

Limitations

  • The study was "confined to freely accessible versions of each AI tool, excluding subscription models that may have more advanced retrieval capabilities and was taken at a specific point in time." Additionally, "while the test set was broad, it remains a small subset of the retraction landscape and may not generalize to all fields, especially since the majority of articles (13/15, 86.7%) were medical articles." The authors also note "the binary scoring approach (correct vs incorrect) captures factual accuracy but does not account for nuances such as ambiguous phrasing." Furthermore, they acknowledge that "our way of questioning needed a uniform, comparable way
  • This was achieved by combining the keywords in a single prompt string, which could differ from a prompt by a common user searching for a specific topic."

Open questions raised

  • Need for expanded sample sizes of retracted articles beyond the 15 tested
  • Evaluation of subscription-based AI models with potentially more advanced retrieval capabilities
  • Exploration of automated cross-validation with bibliometric databases
  • Testing across non-biomedical scientific fields with differing citation and retraction practices
  • Systematic benchmarking and transparent reporting of factual reliability for GenAI systems
  • Development of retraction-aware pipelines with dynamic metadata synchronization
Data: Retraction Watch database (https://retractionwatch.com/); Table S1 in Multimedia Appendix 1 contains details of included retracted articles; Retraction-watch database (https://retraction-watch.pubmed.org/) - for articles retracted closest to May 23, 2025Extracted from: pdfAgreement 57%

Explore related topics

Related papers