Evaluating the predictive capacity of ChatGPT for academic peer review outcomes across multiple platforms
Mike Thelwall, Abdallah Yaghi · Scientometrics · 2025
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1007/s11192-025-05287-1
Methodology & findings
Study design
Multi-case empirical study using correlational analysis.
Sample
N = 750, 9 groups
Primary method
Spearman's rank correlation coefficient was the primary statistical method used to assess the relationship between ChatGPT-averaged scores and human reviewer scores. Bootstrapping in R was used to calculate 95% confidence intervals for population correlations. Custom programming was developed to extract ChatGPT recommendations from text outputs. Two independent programs (in different languages) verified results to rule out programming errors.
Main result
The study found that "Averaging 30 ChatGPT predictions, based on reviewer guidelines and using only the submitted titles and abstracts failed to predict peer review outcomes for F1000Research (Spearman's rho = 0.00). However, it produced mostly weak positive correlations with the quality dimensions of SciPost Physics (rho = 0.25 for validity, rho = 0.25 for originality, rho = 0.20 for significance, and rho = 0.08 for clarity) and a moderate positive correlation for papers from the International Conference on Learning Representations (ICLR) (rho = 0.38)." Including full texts increased the ICLR correlation to rho = 0.46, while chain-of-thought prompting generally decreased effectiveness.
Reports effect sizes and confidence intervals.
Research paradigm
positivist/empiricist
Author conclusions
The authors conclude that "the results suggest that in some contexts, ChatGPT can produce weak pre-publication quality predictions. However, their effectiveness and the optimal strategies for employing them vary considerably between platforms, journals, and conferences. Finally, the most suitable inputs for ChatGPT appear to differ depending on the platform." They further note that "the 95% confidence intervals for population correlations mostly exclude zero (Fig. 5), indicating that ChatGPT averages are likely to be effective for ICLR2017 and SciPost Physics overall. However, the findings suggest that ChatGPT may lack the ability to reliably assess research quality for F1000Research."
Risk of bias
Training data contamination: ChatGPT may have been trained on public peer review data, creating potential memory effects; Selection bias: Only three large platforms with open peer review data were selected; results may not generalize to other journals/conferences; Assumption validity: Analysis assumes human reviewer scores represent 'correct' evaluations, though expert disagreement is common; Platform-specific effects: F1000Research, ICLR, and SciPost Physics have different review formats, discipline foci, and scoring systems, limiting cross-platform comparability; Equation handling bias: Mathematical equations in SciPost Physics papers were reduced to symbolic representations, potentially limiting ChatGPT's ability to assess validity; Training data contamination risk: public review scores may be in ChatGPT training data; Selection bias: only three platforms with publicly available peer review data; Platform non-representativeness: F1000Research multidisciplinary but not representative of all fields; Assumption of human reviewer accuracy: expert reviewers disagree and make mistakes; Outcome measurement bias: conversion of categorical F1000Research decisions (Approve/Approve with Reservations/Not Approved) to numerical scale (1/0.5/0) is arbitrary; Training data contamination risk - ChatGPT may have encountered public peer review scores during training; Selection bias - only three platforms examined; results may not generalize to other fields or publication types; Measurement validity - human reviewers may disagree sharply, making the ground truth assumption unsafe; Platform-specific bias - different review systems and disciplines may have different ChatGPT performance
Limitations
- The authors acknowledge that "since all scores are in the public domain, it is unclear whether ChatGPT had encountered the scores as part of its training data and, if so, whether it retained any useful memory of them
- This is a major conceptual limitation." Additionally, "only three cases were examined, meaning the results may differ substantially for other fields or publication types (e.g., monographs)." Furthermore, "a more fundamental issue is that the analyses implicitly assume that the human reviewers tended to provide the correct recommendations
- Since even expert reviewers disagree, sometimes sharply, this is an unsafe assumption."
Open questions raised
- Whether dataset specificity or review format plays the more important role in ChatGPT's varying performance across platforms
- Fine-grained field-specific analysis (as opposed to multidisciplinary platforms like F1000Research)
- Investigation of alternative prompting strategies that could produce better results
- Testing with competing LLM systems to establish comparative performance
- Extension to other publication types beyond journal articles and conference papers (e.g., monographs)
- Unclear whether dataset-specificity or review format plays the more important role in ChatGPT's varying performance
Explore related topics
Related papers
- Comparing scientific abstracts generated by ChatGPT to real abstracts with detectors and blinded human reviewersCatherine A. Gao · 2023 · 657 citations
- Fabrication and errors in the bibliographic citations generated by ChatGPTWilliam H. Walters · 2023 · 352 citations
- Generative AI tools and assessment: Guidelines of the world's top-ranking universitiesBenjamin Luke Moorhouse · 2023 · 343 citations
- Human-AI collaboration patterns in AI-assisted academic writingAndy Nguyen · 2024 · 301 citations
- Hallucination Rates and Reference Accuracy of ChatGPT and Bard for Systematic Reviews: Comparative AnalysisMikaël Chelli · 2024 · 295 citations
- AI literacy and its implications for prompt engineering strategiesNils Knoth · 2024 · 277 citations