12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

AI-assisted peer review

Alessandro Checco, Lorenzo Bracciale, Pierpaolo Loreti, Stephen Pinfield, Giuseppe Bianchi · Humanities and Social Sciences Communications · 2021

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
2/4
Quality (LMQS)
E
Evidence
261
Citations
10.34
FWCI
Top 10%
Impact

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1057/s41599-020-00703-8

Methodology & findings

Study design

Empirical machine-learning experiment.

Sample

N = 3300, 6 groups

Primary method

Machine learning: Dense neural networks with Rectified Linear Unit (ReLU) activation, dropout regularization, sigmoid output layer. Loss functions: binary cross-entropy (WCNC classification), Mean Squared Error (OpenReview regression). Optimization: Stochastic Gradient Descent (SGD) with Nesterov momentum updates. Model evaluation: train/test split with accuracy, precision, recall, F1-score (WCNC); Mean Absolute Error and Mean Squared Error (OpenReview). Explainability: LIME (Local Interpretable Model-Agnostic Explanations) with submodular pick technique. Baselines: random classifier (WCNC), naive regressor using median score (OpenReview).

Main result

The study found that "despite the focusing on rather superficial document features, like word distribution, readability and formatting scores, the machine-learning system was often able to successfully predict the peer review outcome." For the WCNC dataset, the dense neural network achieved 74.01% accuracy and 72.30% F1-score compared to a random baseline of ~50%. For OpenReview, "75% of the samples have an error of under 1.2 (over a total score of 10), and the median error is only 0.79," demonstrating a strong correlation between superficial metrics and peer review decisions.

Reports effect sizes.

Research paradigm

Empirical, mixed-methods (quantitative machine learning + qualitative analysis)

Author conclusions

The authors conclude: "In this paper, we have reported an experiment involving three peer-reviewed conference proceedings, training a machine-learning system to infer a set of rules able to match the peer review outcome, ultimately providing an acceptance probability for other manuscripts. We focused on a rather superficial set of features of the submitted manuscripts, like word distribution, readability scores and document format. Nevertheless, the machine-learning system was often able to successfully predict the peer review outcome: we found a strong correlation between such superficial features and the outcome of the review process as a whole." They further state that "our findings point in the direction of a new type of analysis of typical human process, conducted with the help of machine-learning systems, one which is cognisant of ethical dimensions of the work as well as technical capabilities."

Risk of bias

Training data bias: Historical peer review decisions may encode existing sociocultural biases in the academic system; Algorithmic bias: AI systems trained on past data may propagate gender, language, and institutional affiliation biases; Geographic bias: Papers from low-income countries or underrepresented regions may be disadvantaged; Overfitting: Authors observed overfitting issues, such as specific first names being regarded as positive signals; Feature bias: Readability and formatting metrics may disproportionately disadvantage non-native English speakers; Design path bias: Models reflect designer values and goals frozen into code; Training bias: Models trained on past peer review decisions may replicate existing biases in historical review data; Geographic/institutional bias: Papers from under-represented regions or institutions may be disadvantaged; Language bias: Non-native English speakers may be penalized on readability metrics; Field-specific bias: Overfitting to conference-specific terminology (e.g., 'decoding' in machine learning conferences); First-impression bias: Superficial document features (formatting, readability) may unduly influence predictions; Selection bias in data collection: Data obtained from two sources (WCNC and OpenReview) with different review processes; Discrepancy between editorial decision and reviewer scores: Authors note "in a small number of cases, there is a rather strong discrepancy between the editorial decision and the average score of the reviewers"; Overfitting: Early stages showed overfitting with specific reviewer names treated as positive signals; Selection bias: Datasets limited to conference proceedings, primarily WCNC (wireless communications) and ICLR (machine learning); results may not generalize to journal submissions or other disciplines; Training data bias: Machine learning models trained on historical review decisions may perpetuate existing biases in the peer review process; Confounding: Papers with better formatting/readability may also have better scientific content; superficial features may be proxies for quality rather than causal drivers; Temporal bias: Models trained on past reviews may not reflect evolving submission quality from historically underrepresented regions; Geographic bias: Authors note USA dominates peer review (32.9% of reviews vs 25.4% of output); model may penalize papers from low-income countries; Overfitting risks: Specific keywords and author names initially flagged as predictive; document-dependent non-linear effects difficult to isolate

Limitations

  • The approach does not cover all aspects of the peer review process, nor does it attempt to replace human reviewers with AI
  • "The approach we take does not cover all the aspects of the peer review process, nor does it attempt to replace human reviewers with AI
  • However, it suggests that there are some components of the quality assessment and peer review process, which could reasonably to assisted or replaced by AI-assisted tools." The authors note that "the possible model of quality assessment we are exploring is then a semi-automated one, where AI informs decision-making, rather than alone determining outcomes
  • It is acknowledged the extent to which this is possible will vary depending on a number of factors, not least the nature of the research output itself and its approach to presenting research results." Additionally, the study does not cover complex intellectual tasks requiring domain expertise, and disciplinary variation may limit generalizability.

Open questions raised

  • Feedback loop: Controlled experiments with academic reviewers to understand biases introduced by AI signals
  • Review process: Incorporating full text of reviews and rebuttals (not just scores) to train AI tools
  • Perception: Extended work on first-impression bias with more complex typographic layout indicators
  • Disciplinary variation: How AI tools apply across different disciplinary norms
  • Grant applications: Application of methods to grant review processes, which have different structure than research manuscripts
  • Feedback loop: "We are interested in exploring the behaviour of reviewers when using these AI-powered support tools. We intend in future to carry out controlled experiments with academic reviewers, to understand the biases introduced by the AI signals on the reviewers."
Data: IEEE WCNC 2018: Available from conference general chair (de-identified); ICLR 2018 and 2019: Available on https://openreview.net/; Generated datasets: Not publicly available due to de-anonymization risks, but available from corresponding author on reasonable request; The paper states: "The datasets generated during the current study are not publicly available due to de-anonymisation risks, but are available from the corresponding author on reasonable request. The OpenReview data are available on https://openreview.net/."; WCNC 2018: 2300 papers with editorial decisions and reviewer scores (internal access only; data de-anonymized); ICLR 2018-2019: 1000 papers from OpenReview (publicly available at https://openreview.net/); Internal datasets not publicly available due to de-anonymisation risks but available from corresponding author on reasonable requestExtracted from: pdfAgreement 55%

Explore related topics

Related papers