12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Evaluating the predictive capacity of ChatGPT for academic peer review outcomes across multiple platforms

Mike Thelwall, Abdallah Yaghi · Scientometrics · 2025

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
3/4
Quality (LMQS)
E
Evidence
17
Citations
7.22
FWCI
Top 10%
Impact

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1007/s11192-025-05287-1

Methodology & findings

Study design

Multi-case empirical study using correlational analysis.

Sample

N = 750, 9 groups

Primary method

Spearman's rank correlation coefficient was the primary statistical method used to assess the relationship between ChatGPT-averaged scores and human reviewer scores. Bootstrapping in R was used to calculate 95% confidence intervals for population correlations. Custom programming was developed to extract ChatGPT recommendations from text outputs. Two independent programs (in different languages) verified results to rule out programming errors.

Main result

The study found that "Averaging 30 ChatGPT predictions, based on reviewer guidelines and using only the submitted titles and abstracts failed to predict peer review outcomes for F1000Research (Spearman's rho = 0.00). However, it produced mostly weak positive correlations with the quality dimensions of SciPost Physics (rho = 0.25 for validity, rho = 0.25 for originality, rho = 0.20 for significance, and rho = 0.08 for clarity) and a moderate positive correlation for papers from the International Conference on Learning Representations (ICLR) (rho = 0.38)." Including full texts increased the ICLR correlation to rho = 0.46, while chain-of-thought prompting generally decreased effectiveness.

Reports effect sizes and confidence intervals.

Research paradigm

positivist/empiricist

Author conclusions

The authors conclude that "the results suggest that in some contexts, ChatGPT can produce weak pre-publication quality predictions. However, their effectiveness and the optimal strategies for employing them vary considerably between platforms, journals, and conferences. Finally, the most suitable inputs for ChatGPT appear to differ depending on the platform." They further note that "the 95% confidence intervals for population correlations mostly exclude zero (Fig. 5), indicating that ChatGPT averages are likely to be effective for ICLR2017 and SciPost Physics overall. However, the findings suggest that ChatGPT may lack the ability to reliably assess research quality for F1000Research."

Risk of bias

Training data contamination: ChatGPT may have been trained on public peer review data, creating potential memory effects; Selection bias: Only three large platforms with open peer review data were selected; results may not generalize to other journals/conferences; Assumption validity: Analysis assumes human reviewer scores represent 'correct' evaluations, though expert disagreement is common; Platform-specific effects: F1000Research, ICLR, and SciPost Physics have different review formats, discipline foci, and scoring systems, limiting cross-platform comparability; Equation handling bias: Mathematical equations in SciPost Physics papers were reduced to symbolic representations, potentially limiting ChatGPT's ability to assess validity; Training data contamination risk: public review scores may be in ChatGPT training data; Selection bias: only three platforms with publicly available peer review data; Platform non-representativeness: F1000Research multidisciplinary but not representative of all fields; Assumption of human reviewer accuracy: expert reviewers disagree and make mistakes; Outcome measurement bias: conversion of categorical F1000Research decisions (Approve/Approve with Reservations/Not Approved) to numerical scale (1/0.5/0) is arbitrary; Training data contamination risk - ChatGPT may have encountered public peer review scores during training; Selection bias - only three platforms examined; results may not generalize to other fields or publication types; Measurement validity - human reviewers may disagree sharply, making the ground truth assumption unsafe; Platform-specific bias - different review systems and disciplines may have different ChatGPT performance

Limitations

  • The authors acknowledge that "since all scores are in the public domain, it is unclear whether ChatGPT had encountered the scores as part of its training data and, if so, whether it retained any useful memory of them
  • This is a major conceptual limitation." Additionally, "only three cases were examined, meaning the results may differ substantially for other fields or publication types (e.g., monographs)." Furthermore, "a more fundamental issue is that the analyses implicitly assume that the human reviewers tended to provide the correct recommendations
  • Since even expert reviewers disagree, sometimes sharply, this is an unsafe assumption."

Open questions raised

  • Whether dataset specificity or review format plays the more important role in ChatGPT's varying performance across platforms
  • Fine-grained field-specific analysis (as opposed to multidisciplinary platforms like F1000Research)
  • Investigation of alternative prompting strategies that could produce better results
  • Testing with competing LLM systems to establish comparative performance
  • Extension to other publication types beyond journal articles and conference papers (e.g., monographs)
  • Unclear whether dataset-specificity or review format plays the more important role in ChatGPT's varying performance
Data: F1000Research articles and reviews (publicly available via f1000research.com); ICLR 2017 metadata and reviewer scores (available from PeerRead repository curated by Kang et al., 2018); SciPost Physics articles and reviews (publicly available via scipost.org); F1000Research papers: crawled from f1000research.com on 7-8 July 2024; ICLR2017 metadata: extracted from repository curated by Kang et al. (2018); SciPost Physics papers: crawled from scipost.org on 8 July 2024, full texts from arXiv or SciPost; F1000Research (public platform, crawled 7-8 July 2024); ICLR2017 metadata repository (Kang et al., 2018 PeerRead dataset); SciPost Physics (crawled 8 July 2024)Code: Webometric Analyst (https://github.com/MikeT​helwa​ll/Webom​etric_​Analy​st) - used for extracting ChatGPT recommendations from reports; Webometric Analyst (custom program for extracting ChatGPT recommendations): https://github.com/MikeThelwall/Webometric_Analyst; Webometric Analyst (https://github.com/MikeT​helwa​ll/Webom​etric_​Analy​st) - custom program for extracting ChatGPT recommendationsExtracted from: pdfAgreement 51%

Explore related topics

Related papers