12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Can social media provide early warning of retraction? Evidence from critical tweets identified by human annotation and large language models

Hui‐Zhen Fu, Mike Thelwall, Zhichao Fang · Journal of the Association for Information Science and Technology · 2025

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
1/4
Quality (LMQS)
E
Evidence
2
Citations
3.53
FWCI
Top 10%
Impact

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1002/asi.70028

Methodology & findings

Study design

Mixed-methods observational study combining human annotation and large language model (LLM) classification.

Sample

N = 7188, 4 groups

Primary method

Coarsened exact matching (CEM) for constructing comparable groups. Gwet's AC1 coefficient for inter-coder reliability assessment. Precision, recall, and F1-Score metrics for LLM classification evaluation. Majority voting procedure across three independent runs per LLM model. Qualitative analysis of false positives. Survival function analysis (cumulative distribution plotting).

Main result

Manual analysis revealed that "8.3% of retracted articles were associated with at least one critical tweet prior to retraction, compared to just 1.5% of non-retracted articles." Additionally, the study found that "3.6% of retracted articles received two or more critical tweets, and 1.0% received ten or more. In contrast, only 0.6% of non-retracted articles received at least two critical tweets, and just 0.1% received ten or more," indicating that retracted articles not only attract more critical tweets overall but also accumulate greater volumes of critical commentary at the individual article level.

Reports effect sizes.

Research paradigm

Positivist/empiricist with computational methods

Author conclusions

"This study examined whether critical tweets can serve as early warning signals for retracted scholarly articles and assessed the potential of LLMs to automate the detection of such tweets. Manual analysis revealed that 8.3% of retracted articles were associated with at least one critical tweet prior to retraction, compared to just 1.5% of non-retracted articles." The authors conclude that "Twitter's potential as a supplementary channel for post-publication review and research integrity monitoring" is supported, but "LLMs' ability to detect critical content aligned only moderately with human annotation" and "are not yet reliable as standalone tools for critical tweet detection. However, they hold promise as components of hybrid workflows that combine automated filtering with human validation to support scalable monitoring."

Risk of bias

Selection bias: Only 29.8% of retracted articles (712 of 2,387) had pre-retraction tweets, potentially biasing toward more visible/discussed articles; Annotation bias: Human judgment is subjective; though inter-coder reliability was high (Gwet's AC1=0.969), residual disagreements were resolved through discussion, potentially introducing consensus bias; Confounding by article topic: No control for thematic content, which likely affects Twitter engagement patterns; Coverage bias: Analysis focused on Twitter only; not representative of broader social media landscape; Temporal bias: Data collected from November 2022 Altmetric snapshot may not capture all pre-retraction tweets; LLM bias: Models trained on heterogeneous corpora may embed social, cultural, or epistemic biases affecting classification accuracy; Selection bias: Only 29.8% of initial retracted articles had pre-retraction tweets, potentially biasing toward more visible articles; Annotation bias: Human annotation relied on subjective judgment despite high inter-coder reliability (Gwet's AC1 = 0.969); Temporal bias: Analysis focused on 2019 publications only to avoid COVID-19 anomalies, limiting generalizability; Ideological bias: Critical tweets may reflect disagreement rather than actual methodological flaws; Platform bias: Analysis limited to Twitter, excluding other social media platforms; LLM biases: Models may embed social, cultural, or epistemic biases from training corpora; Selection bias: Only articles with pre-retraction tweets included (29.8% of initial sample of 2387 retracted articles); Annotation bias: Human judgment inherently subjective despite high inter-coder agreement (Gwet's AC1 = 0.969); LLM bias: Models may embed social, cultural, or epistemic biases from training corpora; Platform bias: Twitter data only; not generalizable to other social media platforms; Temporal bias: COVID-19 articles excluded due to anomalous retraction patterns; User representation bias: Analysis did not account for bot accounts, which play notable role in disseminating scientific content

Limitations

  • The authors identified several key limitations: "First, although human annotation was used as the benchmark for evaluating LLM performance, human judgment is inherently subjective and susceptible to bias." Second, "not all critical tweets reflect verifiable issues in the referenced articles, nor do they reliably predict eventual retractions
  • In some cases, genuine criticism may arise from ideological disagreement rather than methodological flaws." Third, "this study did not differentiate between retraction reasons, such as methodological errors, data fabrication, ethical violations, or plagiarism." Fourth, "the analysis focused exclusively on tweet text and did not incorporate engagement metrics (e.g., likes, retweets, and replies) or contextual data about users." Finally, "recent structural changes to Twitter—such as the shift to paid API access and stricter rate limits—have constrained researchers' ability to collect large-scale, up-to-date datasets."

Open questions raised

  • Future research should investigate whether specific categories of retraction (methodological errors, data fabrication, ethical violations, plagiarism) are more readily detectable through critical online discourse
  • Thematic content analysis needed to clarify relationship between article content and critique patterns
  • Incorporation of engagement metrics (likes, retweets, replies) and user contextual data (bot vs. human accounts)
  • Extension to emerging platforms (Bluesky, Mastodon) as Twitter's representation diminishes
  • Integration of additional contextual signals to improve detection accuracy
  • Differentiation between retraction reasons (methodological errors, data fabrication, ethical violations, plagiarism) and their distinct patterns of social media engagement
Data: Web of Science Core Collection snapshot (March 2025) - restricted access via CWTS at Leiden University; Retraction Watch database (accessed June 2025) - publicly available; Altmetric database snapshot (November 2022) - restricted access via CWTS at Leiden University; Tweet data retrieved via Twitter API (March 2023) - subject to Twitter/X API terms and rate limits; Web of Science Core Collection (snapshot dated March 2025); Altmetric database (snapshot dated November 2022); Retraction Watch database (accessed June 2025); Web of Science (WoS) database snapshot (March 2025) - maintained by CWTS at Leiden University; Altmetric database snapshot (November 2022) - maintained by CWTS; Tweet data retrieved via Twitter API (March 2023) - publicly available but subject to Twitter API termsExtracted from: pdfAgreement 51%

Explore related topics

Related papers