Human-AI Collaboration in Science at Scale: A Global Large-scale Randomized Field Experiment
Binglu Wang, Weixin Liang, Jiahui Xue, Yuhui Zhang, Hancheng Cao, Dashun Wang et al. · arXiv (Cornell University) · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
Methodology & findings
Study design
Global large-scale randomized controlled field experiment.
Sample
N = 45466, 6 groups
Primary method
OLS regression models at paper level (revisions outcome) and author level (adoption outcome). Stratified randomization by field. Intent-to-treat analysis. Robustness checks with alternative post-treatment windows. Computational pipeline for content analysis comparing original and revised manuscripts. State-of-the-art AI detection model (Liang et al. 2025) for LLM adoption measurement. Software/packages not explicitly specified.
Main result
The study found that "authors who received feedback had a significantly higher likelihood of revising their manuscripts, corresponding to a 12.55% relative increase over the baseline revision rate." Additionally, "exposure to AI feedback also increased authors' subsequent use of LLM tools in their future papers, suggesting longer-run shifts in scientific practice." The effects were "strongest among authors from non-English-dominant research regions, manuscripts less embedded in the scholarly literature, and teams with lower h-indexes and earlier career stages."
Reports effect sizes and confidence intervals.
Research paradigm
Positivist/empirical-analytical; quantitative causal inference
Author conclusions
"This study provides causal evidence that large language models can meaningfully participate in a central collaborative practice of science: offering feedback. In a global large-scale field experiment, exposure to AI-generated critique prompted substantive revisions to manuscripts and increased subsequent adoption of AI tools. These effects were not evenly distributed. They were stronger among manuscripts and authors positioned further from established linguistic, intellectual, or professional advantage, suggesting that AI has the potential to act as a collaborative equalizer in science, expanding access to critique where traditional channels are less accessible."
Risk of bias
Attrition: 3,320 manuscripts (9.7%) posted updated versions during feedback generation period, excluded from analysis; Detection bias: LLM-detection methods inherently noisy and may show differential accuracy for non-native English writing; Hawthorne effect: Treatment effect may partially reflect salience of receiving correspondence rather than feedback content alone; Intent-to-treat bias: Estimates include authors who did not open or engage with feedback, representing conservative lower bound; Attention/salience bias: Control group received no contact, making it difficult to disentangle feedback content effects from mere attention effects; Measurement bias: LLM detection models may exhibit differential accuracy across non-native English writing styles; Attrition: 9.7% of manuscripts (3,320) posted updated versions between feedback generation and delivery and were excluded from analysis; Intent-to-treat bias: Estimates include authors who did not open or engage with feedback, potentially underestimating complier effects; Detection model bias: Inherent noisiness in AI detection methodology may differentially affect treatment and control groups; No-contact control group design cannot distinguish feedback content effects from attention/salience effects; LLM detection model for adoption outcome is inherently noisy with potential differential accuracy across groups (particularly non-native English writers); Intent-to-treat analysis includes non-compliers who did not access feedback, potentially underestimating effects; Attrition: 9.7% of manuscripts (3,320) posted updated versions between feedback generation and delivery, excluded from analysis; Self-selected adoption of arXiv/email contact; authors must check email and access private webpage; Potential differential engagement based on language proficiency and digital literacy not directly measured
Limitations
- The study notes that "Our design compares treated authors to a no-contact control group, raising the question of whether our results reflect feedback content or the attentional salience of receiving any correspondence" and that "Our adoption measure relies on LLM-detection methods, which are inherently noisy and may exhibit differential accuracy, such as for non-native English writing styles." The authors also state that "our estimates reflect intent-to-treat effects from a single light-touch intervention, necessarily including authors who did not open or engage with the feedback
- they are therefore best interpreted as a conservative lower bound on the effect among active compliers."
Open questions raised
- Authors identify the need for future work to: (1) continue triangulating LLM adoption measures with ground-truth data on actual tool use; (2) understand how AI reshapes inequalities in knowledge production; (3) preserve heterogeneity in human-AI collaboration through 'plural model architectures, varied objectives, and institutional norms that reward dissent and methodological diversity' to counteract standardization and homogenization risks in scientific feedback.
- Future work should triangulate LLM detection measures with ground-truth data on actual tool use
- Need to explore how to preserve heterogeneity in AI-assisted science to avoid homogenization of scientific feedback
- Investigation of whether human collaborators shift toward higher-order conceptual and design roles as AI provides routine critique
- Study of institutional norms that reward dissent and methodological diversity in light of standardized AI feedback
- The authors identify several gaps and future directions: (1) need for continued triangulation of LLM detection measures with ground-truth data on actual tool use; (2) designing human-AI collaboration to preserve heterogeneity through plural model architectures and varied objectives to counter potential homogenization of scientific recommendations; (3) importance of institutional norms that reward dissent and methodological diversity; (4) ongoing investigation of how AI reshapes collaboration as technology evolves and adoption accelerates beyond baseline use.
Explore related topics
Related papers
- Estimating the reproducibility of psychological scienceAlexander A. Aarts · 2015 · 8,669 citations
- The effects of over-reliance on AI dialogue systems on students' cognitive abilities: a systematic reviewChunpeng Zhai · 2024 · 1,009 citations
- Autonomous chemical research with large language modelsDaniil A. Boiko · 2023 · 809 citations
- To use or not to use ChatGPT in higher education? A study of students’ acceptance and use of technologyArtur Strzelecki · 2023 · 691 citations
- Teacher support and student motivation to learn with Artificial Intelligence (AI) based chatbotThomas K. F. Chiu · 2023 · 617 citations
- Do AI chatbots improve students learning outcomes? Evidence from a meta‐analysisRong Wu · 2023 · 469 citations