When AI Meets Science: Research Diversity, Interdisciplinarity, Visibility, and Retractions across Disciplines in a Global Surge
Andrés F. Castro Torres, Joan Giner-Miguelez, Mercè Crosas · ArXiv.org · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
Methodology & findings
Study design
Large-scale bibliometric analysis using the OpenAlex collection (227 million scholarly works from 1960-2024) spanning four scientific domains and 46 fields.
Sample
N = 227854643, 8 groups
Primary method
Dictionary-based word search methods with tokenization and term tagging. Large Language Model (LLM) semantic classification using Qwen2.5-7B-Instruct deployed on HPC infrastructure (92 hours compute across 96 NVIDIA A100 GPUs). Retrieval-Augmented Generation (RAG) for methods section identification in full texts using all-MiniLM-L6-v2 sentence-transformers. Inter-rater agreement assessed via Cohen's Kappa and Krippendorff's Alpha. Logistic regression models predicting retraction rates and citation patterns, accounting for publication type, language, and publication year. Cosine similarity measures for semantic heterogeneity analysis. Predicted probabilities and marginal effects reported across scientific fields.
Main result
The study found that "AI-supported research is confined to a few topics with strong ties to Computer Science and conventional statistical frameworks, suggesting limited epistemological transformation. It is also associated with an unwarranted citation premium and substantially higher retraction rates than non-AI-supported" research. Additionally, "Geographically, while wealthy countries lead in AI publications per capita, global South countries in a belt from Indonesia to Algeria lead in AI adoption relative to their national output, signaling a distinctive resource concentration pattern."
Reports effect sizes.
Research paradigm
Quantitative empirical analysis using bibliometrics and computational methods
Author conclusions
The authors conclude that "The transformative capacity of AI in science thus remains untapped, and its rapid adoption underlines challenges in research openness, transparency, reproducibility, and ethics." They further state that "Greater attention to these historical accounts could help guide AI adoption to ensure that the principles of openness, transparency, and ethics are integral to the roots of the AI technology tree." Additionally, they emphasize that "Decreased topic diversity, greater reliance on linear model analysis frameworks, and a higher likelihood of citations to Computer Science works (but not to other fields) compared to non-AI-based research suggest that AI adoption's potential is still untapped across most scientific fields."
Risk of bias
Selection bias: Focus on works with abstracts ≥200 characters may exclude certain publication types; Measurement bias: LLM-based classification may miss or misclassify AI methods, particularly those using non-standard terminology; Database bias: OpenAlex collection may have differential coverage across countries, languages, and disciplines; Temporal bias: Conservative estimates for earlier periods due to less standardized AI methodology reporting; Geographic bias: Analysis relies on first author institutional affiliation, potentially obscuring international collaboration; Publication bias: Restriction to peer-reviewed works excludes preprints and grey literature patterns; Detection bias: Abstract-based approach underestimates AI methods compared to full-text analysis (2.1% vs 7.6% in PLOS ONE sample); Publication bias: Analysis only includes published/indexed works in OpenAlex; Language bias: Dictionary expanded to include translations, but non-English publications may be underrepresented; Domain bias: Differential detection gaps across domains (Life Sciences 1.6% vs Physical Sciences 5.0% in abstracts); Coder heterogeneity: Inter-coder agreement metrics (Cohen's Kappa=0.58-0.81) suggest variable classification reliability across coders with different backgrounds; Potential underestimation of AI adoption in abstracts compared to full-text analysis (gap of 1.1-15.4 percentage points depending on domain); Selection bias in abstract-based approach: abstracts may not reflect actual methodological content; False positive/negative errors in LLM classification despite validation (6.7% false negatives in AI methods classification; 0.88% false positives); Language bias: non-English publications may be underrepresented in OpenAlex; Publication bias inherent in bibliometric data: only published/indexed works analyzed; Confounding by publication type, language, and publication year (controlled for in regression analyses); Validation sample (911 abstracts from 2012-2024) may not represent earlier periods (1960-2011); Researcher coding diversity bias: coders had different academic backgrounds and training levels affecting inter-rater agreement
Limitations
- The authors note that "abstracts may underperform in cases where abstracts do not explicitly describe the research methods, methods' descriptions deviate from conventional phrasing, or there are no conventions to refer to methods." Additionally, "Due to the word-search approach, it remains unclear whether this growth and integration of AI into the scientific literature is due to increased AI discussion, adoption, and development." The authors also acknowledge that "Another limitation of these studies is that AI-use prevalence and growth are measured with respect to the entire body of academic works
- This approach mixes articles that may not be susceptible to methods of empirical analysis." Furthermore, they note that "Cohen's Kappa and Krippendorff's Alpha statistics at 0.58 reflected the complexity of identifying methods in abstracts" and that inter-coder agreement was imperfect at 79.1% for methods identification and 75.4% for AI/non-AI classification.
Open questions raised
- Understanding how AI adoption varies across institutional arrangements in different countries
- Clarifying whether higher citation rates for AI-based research reflect genuine quality advantages or represent hype-driven attention patterns
- Investigating structural factors contributing to higher retraction rates in AI-based research
- Developing frameworks for research openness, transparency, and ethics specific to AI technologies and LLM applications
- Understanding patterns of AI adoption in Sub-Saharan Africa and other regions with low adoption rates
- Examining gender and racial/ethnic differentials in AI adoption across different fields and regions
Explore related topics
Related papers
- Practical and ethical challenges of large language models in education: A systematic scoping reviewLixiang Yan · 2023 · 699 citations
- Comparing scientific abstracts generated by ChatGPT to real abstracts with detectors and blinded human reviewersCatherine A. Gao · 2023 · 657 citations
- ChatGPT and the rise of generative AI: Threat to academic integrity?Damian Eke · 2023 · 476 citations
- Fabrication and errors in the bibliographic citations generated by ChatGPTWilliam H. Walters · 2023 · 352 citations
- Human-AI collaboration patterns in AI-assisted academic writingAndy Nguyen · 2024 · 301 citations
- Hallucination Rates and Reference Accuracy of ChatGPT and Bard for Systematic Reviews: Comparative AnalysisMikaël Chelli · 2024 · 295 citations