AI-Augmented Peer Review and Scientific Productivity: A Cross-Country Panel and SEM Analysis
Dongsoo Han · arXiv (Cornell University) · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
Methodology & findings
Study design
Cross-country panel regression analysis (fixed-effects models), mediation analysis with bootstrap confidence intervals (N=1,000 replications), and structural equation modeling (SEM) using OECD country data from an unspecified time period.
Sample
> 1000, 2 groups
Primary method
Fixed-effects panel regression with country and year fixed effects; lagged variable specifications to address simultaneity bias; instrumental variable (IV) approach using early internet adoption and digital infrastructure as instruments; mediation analysis with bootstrap confidence intervals (1,000 replications); structural equation modeling (SEM) with latent variables; dynamic panel GMM estimation; robustness checks using alternative dependent variables and AIRC measures; subsample analyses comparing high-income vs. middle-income countries and pre- vs. post-periods
Main result
The study found that "a one standard deviation increase in AIRC is associated with an [XX]% increase in productivity" and that "the majority of AI's impact on productivity operates through the mediating pathways of review efficiency and reproducibility." The analysis demonstrates that "the indirect effect via review efficiency ([X.XX]) is the largest single pathway, consistent with the hypothesis that AI primarily accelerates the evaluation process, thereby reducing time-to-publication and enabling faster knowledge dissemination."
Reports effect sizes and confidence intervals.
Research paradigm
positivist/empiricist with quantitative econometric analysis
Author conclusions
The authors conclude: "This study provides the first cross-country empirical analysis of the impact of AI-augmented peer review systems on scientific productivity. Using a panel dataset covering OECD countries from [period] to [period], we demonstrate that higher levels of AI Review Capability (AIRC) are associated with significantly higher scientific productivity, with a one standard deviation increase in AIRC corresponding to an 9-12% increase in productivity." They further argue that "The integration of AI into peer review-through hybrid AI-human models that combine computational efficiency with human judgment-represents a promising path toward a more productive, reliable, and equitable scientific system."
Risk of bias
Reverse causality: higher productivity may drive more AI adoption rather than AI driving productivity; Omitted variable bias: institutional quality may affect both AI adoption and productivity; Proxy measurement of AIRC using indirect indicators rather than direct measures; Geographic limitation to OECD countries may introduce selection bias; Temporal ordering challenges despite use of lagged variables; Reverse causality: higher productivity may drive more AI adoption rather than vice versa; Endogeneity concerns acknowledged but only partially addressed through lagged variables and instrumental variable approach; Measurement error in AIRC proxy variables; Limited to OECD countries, potentially missing important global variation; Temporal ordering issues in establishing causality despite use of lagged variables; Measurement error in proxy variables for AIRC at national level; Selection bias: analysis limited to OECD countries only; Endogeneity concerns in panel data
Open questions raised
- Limited evidence on how AI integration affects scientific productivity across countries and over time
- Underexplored mechanisms through which AI influences scientific outcomes, particularly the roles of review efficiency and reproducibility as mediating factors
- Lack of integrated frameworks combining AI, human judgment, and community-based evaluation
- Need for more precise national-level indicators of AI usage in peer review
- Insufficient investigation of ethical and governance issues including bias, transparency, and accountability in AI peer review
- Lack of evidence in non-OECD countries and emerging economies
Explore related topics
Related papers
- Comparing scientific abstracts generated by ChatGPT to real abstracts with detectors and blinded human reviewersCatherine A. Gao · 2023 · 657 citations
- Fabrication and errors in the bibliographic citations generated by ChatGPTWilliam H. Walters · 2023 · 352 citations
- Generative AI tools and assessment: Guidelines of the world's top-ranking universitiesBenjamin Luke Moorhouse · 2023 · 343 citations
- Human-AI collaboration patterns in AI-assisted academic writingAndy Nguyen · 2024 · 301 citations
- Hallucination Rates and Reference Accuracy of ChatGPT and Bard for Systematic Reviews: Comparative AnalysisMikaël Chelli · 2024 · 295 citations
- AI literacy and its implications for prompt engineering strategiesNils Knoth · 2024 · 277 citations