Changes in Manuscript Length, Research Team Size, and International Collaboration in the Post-2022 Period: Evidence from PLOS ONE
Yossi Ben-Zion, Eden Cohen, Nitza Davidovitch · ArXiv.org · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
Methodology & findings
Study design
Bibliometric analysis of a complete population of 109,393 research articles published in PLOS ONE between 2019 and 2025.
Sample
N = 109393, 10 groups
Primary method
One-way ANOVA on log10-transformed word counts with Tukey HSD post-hoc comparisons; two-way ANOVA with year and language background/continent/field as factors; Kruskal–Wallis non-parametric tests (robustness checks); linear regression and difference-in-differences models on log10(word count) with field fixed effects; negative binomial regression with log link for count outcomes (authorship team size, reference counts); logistic regression for binary outcome (collaboration with NES co-authors). Software: R 4.5.0 (base functions, MASS package for negative binomial, emmeans package for post-hoc contrasts), Python (pandas, matplotlib for figures). All analyses on log10-transformed word counts due to right-skewed distributions; complete-case analysis for missing data.
Main result
The study found that "manuscript length increased substantially, with gains ranging from 14.8% among African-affiliated authors and 11.7% among Asian-affiliated authors to 5.3% among native English-speaking (NES) authors, cutting the word-count gap by 39%." Additionally, "non-native English-speaking (NNES) authors reduced both authorship team size, from 6.54 to 6.06 authors, or 7.3%, and collaboration with NES co-authors, from 17.8% to 12.2%, or 36%, while NES authors remained stable in both team size and collaboration rates."
Reports effect sizes and confidence intervals.
Research paradigm
Positivist; quantitative empirical observation
Author conclusions
"This study documents three concurrent post-2022 shifts in scientific publishing: manuscripts grew longer, the word-count gap between NES and NNES authors narrowed, and NNES authors reduced both their team sizes and their collaboration with NES co-authors." The authors conclude that "generative language tools may be reshaping not only how scientific texts are produced, but how collaborative structures are organized, partially substituting for relationships that previously served a linguistic function." They note that "Whether this reorganization represents a net benefit for the global scientific system remains an open question."
Risk of bias
Selection bias: Single journal analysis limits generalizability; Classification bias: LLM-based language background classification relies on institutional affiliation rather than actual linguistic background; potential misclassification of researchers in English-majority institutions versus actual native speakers; Confounding: Policy-level constraints on international collaboration (e.g., China restrictions) not fully quantified; Compositional bias: Disciplinary shifts across time partially controlled via field fixed effects; Temporal confounding: Cannot establish causality between LLM diffusion and observed changes; Selection bias: Sample restricted to single journal (PLOS ONE); findings may not generalize to other journals with different selectivity/editorial policies; Classification bias: Author language background classified by institutional affiliation rather than actual linguistic background; authors in English-majority countries coded as NES even if NNES by birth/first language; Confounding by policy: International collaboration decline may partly reflect China policy constraints rather than LLM adoption; Temporal confounding: Cannot rule out concurrent structural influences on academic publishing beyond LLM diffusion; Compositional bias: Field distribution changes across years partially addressed through fixed effects, but residual concerns remain; Observational design without causal control condition; Misclassification bias: author linguistic background determined by institutional affiliation rather than actual first language; Temporal confounding: unable to isolate effect of LLM diffusion from concurrent structural influences; Selection bias from single journal (PLOS ONE) with consistent editorial framework may not generalize; Unobserved policy confounders: international collaboration restrictions (notably in China) not quantified
Limitations
- "First, the analysis is restricted to a single journal, PLOS ONE
- Although this design offers the methodological advantage of a consistent editorial framework across disciplines and time, it limits the generalizability of the findings." Additionally, "the study identifies temporal associations between post-2022 trends and structural publication characteristics but cannot establish causal attribution
- No direct measure of LLM usage is available at the document level, and no external control condition exists." Furthermore, "the classification of authors as NES or NNES relies on institutional affiliation rather than on the author's actual linguistic background
- This operationalization introduces systematic misclassification."
Open questions raised
- The authors identify several directions for future research: (1) Replication across journals of varying selectivity; (2) Establishing causal attribution through direct measures of LLM usage; (3) Quantifying the extent of policy-driven versus technology-driven effects on international collaboration, particularly in China; (4) Examining whether textual expansion translates into measurable improvements in scientific quality or impact; (5) Investigating mechanisms driving differential adoption across linguistic and disciplinary groups; (6) Continued longitudinal analysis to determine whether documented patterns reflect temporary adjustment or durable structural shift.
- Replication across journals of varying selectivity needed before broader conclusions
- Direct measurement of LLM usage at document level not available; cannot establish causality
- Mechanisms driving differential adoption across linguistic and disciplinary groups require investigation
- Whether textual expansion translates into measurable improvements in scientific quality or impact requires empirical testing
- Disentangling policy-driven collaboration constraints (e.g., China policy) from technology-driven effects
Explore related topics
Related papers
- Autonomous chemical research with large language modelsDaniil A. Boiko · 2023 · 809 citations
- To use or not to use ChatGPT in higher education? A study of students’ acceptance and use of technologyArtur Strzelecki · 2023 · 691 citations
- Comparing scientific abstracts generated by ChatGPT to real abstracts with detectors and blinded human reviewersCatherine A. Gao · 2023 · 657 citations
- Teacher support and student motivation to learn with Artificial Intelligence (AI) based chatbotThomas K. F. Chiu · 2023 · 617 citations
- Do AI chatbots improve students learning outcomes? Evidence from a meta‐analysisRong Wu · 2023 · 469 citations
- Beware of metacognitive laziness: Effects of generative artificial intelligence on learning motivation, processes, and performanceYizhou Fan · 2024 · 419 citations