Can AI Solve the Peer Review Crisis? A Large Scale Cross Model Experiment of LLMs' Performance and Biases in Evaluating over 1000 Economics Papers
Pat Pataranutaporn, Nattavudh Powdthavee · ArXiv.org · 2025
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.48550/arxiv.2502.00070
Methodology & findings
Study design
Large-scale experimental design with systematic variation of author characteristics.
Sample
N = 9030, 6 groups
Primary method
Ordinary Least Squares (OLS) regression with bootstrap standard errors (1,000 replications); ordered logit models with bootstrap standard errors for ordinal outcomes; submission fixed effects analysis; Westfall and Young (1993) free step-down resampling method for controlling family-wise error rate (FWER) across multiple comparisons; predicted margins from logit models.
Main result
The study found that "LLM is highly effective at distinguishing between submissions published in low-, medium-, and high-quality journals. This result highlights the LLM's potential to reduce editorial workload and expedite the initial screening process significantly. However, it struggles to differentiate high-quality papers from AI-generated submissions crafted to resemble "top five" journal standards." Additionally, "there is compelling evidence of a modest but consistent premium-approximately 2-3%-associated with papers authored by prominent individuals, male economists, or those affiliated with elite institutions compared to blind submissions."
Reports effect sizes and confidence intervals.
Research paradigm
Positivist/empiricist (quantitative experimental design with causal inference)
Author conclusions
"By utilizing an experimental design commonly employed in lab and field settings, it highlights the strengths and weaknesses of LLMs in evaluating economics papers, offering a solid foundation for future research and practical improvements in the peer review process. These findings carry critical implications for both the efficiency and equity of academic publishing. On the one hand, the LLM's strong performance in distinguishing paper quality suggests that AI has considerable potential to streamline editorial workflows, especially in the early stages of desk rejection. On the other hand, its susceptibility to biases and inability to detect unethical practices highlights the need for cautious integration." Furthermore, "by refining AI algorithms to prioritize intrinsic paper quality over author attributes, implementing post-hoc adjustments to mitigate biases, and adopting hybrid review models that integrate human judgment with AI evaluations, journals can leverage the advantages of AI while minimizing its risks to fairness in the peer review process."
Risk of bias
Selection bias in paper choice (only recently published papers from 2024-2025); Potential confounding between institutional affiliation and paper quality; Limited generalizability to non-economics disciplines; Unmeasured factors: race, ethnicity, geographic location of institutions; AI training data may encode historical biases from prior publications; Selection bias: Papers were from published sources only, not submissions; Limited demographic variation: Study focused on gender, affiliation, and prominence but excluded race and geographic biases; Confounding: Cannot disentangle whether institutional affiliation reflects genuine quality differences or pure bias effects; Artificial task setting: LLM evaluations in controlled setting may not reflect real editorial decisions with time pressure and reviewer fatigue; Data contamination: Papers from 2024-2025 selected to avoid LLM prior knowledge, but some papers from Asian Economic and Financial Review appeared online in 2022; Selection bias: Papers were intentionally selected from recently published works (2024-2025) to prevent prior LLM knowledge, which may not be representative of typical submission pools; Confounding: Affiliation, prominence, and gender correlate in real-world academic publishing; the orthogonal experimental design does not capture these natural correlations; Limited generalizability: Study uses only 30 base papers; results may not generalize across different paper types, methodologies, or disciplines; AI model-specific findings: Results limited to GPT4o-mini; other LLMs may exhibit different bias patterns; Lack of within-submission comparisons for blind submissions: Only 30 blind submissions versus 9,000 non-blind submissions creates imbalance; Omitted variables: Race/ethnicity, geographic location, and other demographic factors not experimentally varied
Limitations
- "One limitation of our findings is that we cannot directly compare the effect size of AI biases with human biases
- While previous research provides some insights into the magnitude of human biases toward prominent authors (Huber et al., 2022), our study is not directly comparable, as effect sizes are likely context-dependent." Furthermore, "the experimental design systematically varied author characteristics, such as institutional affiliation, reputation, and gender, but did not account for other potentially significant factors, such as race-which may be inferred from author names (Bertrand & Mullainathan, 2004)-or the geographic location of the author's institution." Additionally, "Another limitation of our results is that the sampled paper submissions were drawn from already published papers, suggesting that the actual acceptance rate for unpublished papers may be even lower than the numbers reported in this study."
Open questions raised
- Need to investigate whether AI systems treat identical papers differently based on author characteristics (previously unexplored with randomization)
- Limited research directly comparing AI-generated ratings with human judgments of paper quality using already published papers
- Missing investigation of racial and regional biases that may influence both AI and human evaluations
- Lack of comparative analysis between AI bias magnitudes and human reviewer biases in real editorial decisions
- Need for studies examining the amplification of biases by AI systems in practice
- Whether AI systems treat identical papers differently based on authors' affiliations, gender, or prominence - addressed in this study through experimental approach
Explore related topics
Related papers
- Estimating the reproducibility of psychological scienceAlexander A. Aarts · 2015 · 8,669 citations
- Autonomous chemical research with large language modelsDaniil A. Boiko · 2023 · 809 citations
- Teacher support and student motivation to learn with Artificial Intelligence (AI) based chatbotThomas K. F. Chiu · 2023 · 617 citations
- Beware of metacognitive laziness: Effects of generative artificial intelligence on learning motivation, processes, and performanceYizhou Fan · 2024 · 419 citations
- Impact of AI assistance on student agencyAli Darvishi · 2023 · 384 citations
- Generative AI tools and assessment: Guidelines of the world's top-ranking universitiesBenjamin Luke Moorhouse · 2023 · 343 citations