Can social media provide early warning of retraction? Evidence from critical tweets identified by human annotation and large language models
Hui‐Zhen Fu, Mike Thelwall, Zhichao Fang · Journal of the Association for Information Science and Technology · 2025
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1002/asi.70028
Methodology & findings
Study design
Mixed-methods observational study combining human annotation and large language model (LLM) classification.
Sample
N = 7188, 4 groups
Primary method
Coarsened exact matching (CEM) for constructing comparable groups. Gwet's AC1 coefficient for inter-coder reliability assessment. Precision, recall, and F1-Score metrics for LLM classification evaluation. Majority voting procedure across three independent runs per LLM model. Qualitative analysis of false positives. Survival function analysis (cumulative distribution plotting).
Main result
Manual analysis revealed that "8.3% of retracted articles were associated with at least one critical tweet prior to retraction, compared to just 1.5% of non-retracted articles." Additionally, the study found that "3.6% of retracted articles received two or more critical tweets, and 1.0% received ten or more. In contrast, only 0.6% of non-retracted articles received at least two critical tweets, and just 0.1% received ten or more," indicating that retracted articles not only attract more critical tweets overall but also accumulate greater volumes of critical commentary at the individual article level.
Reports effect sizes.
Research paradigm
Positivist/empiricist with computational methods
Author conclusions
"This study examined whether critical tweets can serve as early warning signals for retracted scholarly articles and assessed the potential of LLMs to automate the detection of such tweets. Manual analysis revealed that 8.3% of retracted articles were associated with at least one critical tweet prior to retraction, compared to just 1.5% of non-retracted articles." The authors conclude that "Twitter's potential as a supplementary channel for post-publication review and research integrity monitoring" is supported, but "LLMs' ability to detect critical content aligned only moderately with human annotation" and "are not yet reliable as standalone tools for critical tweet detection. However, they hold promise as components of hybrid workflows that combine automated filtering with human validation to support scalable monitoring."
Risk of bias
Selection bias: Only 29.8% of retracted articles (712 of 2,387) had pre-retraction tweets, potentially biasing toward more visible/discussed articles; Annotation bias: Human judgment is subjective; though inter-coder reliability was high (Gwet's AC1=0.969), residual disagreements were resolved through discussion, potentially introducing consensus bias; Confounding by article topic: No control for thematic content, which likely affects Twitter engagement patterns; Coverage bias: Analysis focused on Twitter only; not representative of broader social media landscape; Temporal bias: Data collected from November 2022 Altmetric snapshot may not capture all pre-retraction tweets; LLM bias: Models trained on heterogeneous corpora may embed social, cultural, or epistemic biases affecting classification accuracy; Selection bias: Only 29.8% of initial retracted articles had pre-retraction tweets, potentially biasing toward more visible articles; Annotation bias: Human annotation relied on subjective judgment despite high inter-coder reliability (Gwet's AC1 = 0.969); Temporal bias: Analysis focused on 2019 publications only to avoid COVID-19 anomalies, limiting generalizability; Ideological bias: Critical tweets may reflect disagreement rather than actual methodological flaws; Platform bias: Analysis limited to Twitter, excluding other social media platforms; LLM biases: Models may embed social, cultural, or epistemic biases from training corpora; Selection bias: Only articles with pre-retraction tweets included (29.8% of initial sample of 2387 retracted articles); Annotation bias: Human judgment inherently subjective despite high inter-coder agreement (Gwet's AC1 = 0.969); LLM bias: Models may embed social, cultural, or epistemic biases from training corpora; Platform bias: Twitter data only; not generalizable to other social media platforms; Temporal bias: COVID-19 articles excluded due to anomalous retraction patterns; User representation bias: Analysis did not account for bot accounts, which play notable role in disseminating scientific content
Limitations
- The authors identified several key limitations: "First, although human annotation was used as the benchmark for evaluating LLM performance, human judgment is inherently subjective and susceptible to bias." Second, "not all critical tweets reflect verifiable issues in the referenced articles, nor do they reliably predict eventual retractions
- In some cases, genuine criticism may arise from ideological disagreement rather than methodological flaws." Third, "this study did not differentiate between retraction reasons, such as methodological errors, data fabrication, ethical violations, or plagiarism." Fourth, "the analysis focused exclusively on tweet text and did not incorporate engagement metrics (e.g., likes, retweets, and replies) or contextual data about users." Finally, "recent structural changes to Twitter—such as the shift to paid API access and stricter rate limits—have constrained researchers' ability to collect large-scale, up-to-date datasets."
Open questions raised
- Future research should investigate whether specific categories of retraction (methodological errors, data fabrication, ethical violations, plagiarism) are more readily detectable through critical online discourse
- Thematic content analysis needed to clarify relationship between article content and critique patterns
- Incorporation of engagement metrics (likes, retweets, replies) and user contextual data (bot vs. human accounts)
- Extension to emerging platforms (Bluesky, Mastodon) as Twitter's representation diminishes
- Integration of additional contextual signals to improve detection accuracy
- Differentiation between retraction reasons (methodological errors, data fabrication, ethical violations, plagiarism) and their distinct patterns of social media engagement
Explore related topics
Related papers
- Estimating the reproducibility of psychological scienceAlexander A. Aarts · 2015 · 8,669 citations
- Academic Integrity considerations of AI Large Language Models in the post-pandemic era: ChatGPT and beyondMike Perkins · 2023 · 668 citations
- Comparing scientific abstracts generated by ChatGPT to real abstracts with detectors and blinded human reviewersCatherine A. Gao · 2023 · 657 citations
- ChatGPT and the rise of generative AI: Threat to academic integrity?Damian Eke · 2023 · 476 citations
- Nonhuman “Authors” and Implications for the Integrity of Scientific Publication and Medical KnowledgeAnnette Flanagin · 2023 · 399 citations
- Fabrication and errors in the bibliographic citations generated by ChatGPTWilliam H. Walters · 2023 · 352 citations