AI-assisted peer review
Alessandro Checco, Lorenzo Bracciale, Pierpaolo Loreti, Stephen Pinfield, Giuseppe Bianchi · Humanities and Social Sciences Communications · 2021
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1057/s41599-020-00703-8
Methodology & findings
Study design
Empirical machine-learning experiment.
Sample
N = 3300, 6 groups
Primary method
Machine learning: Dense neural networks with Rectified Linear Unit (ReLU) activation, dropout regularization, sigmoid output layer. Loss functions: binary cross-entropy (WCNC classification), Mean Squared Error (OpenReview regression). Optimization: Stochastic Gradient Descent (SGD) with Nesterov momentum updates. Model evaluation: train/test split with accuracy, precision, recall, F1-score (WCNC); Mean Absolute Error and Mean Squared Error (OpenReview). Explainability: LIME (Local Interpretable Model-Agnostic Explanations) with submodular pick technique. Baselines: random classifier (WCNC), naive regressor using median score (OpenReview).
Main result
The study found that "despite the focusing on rather superficial document features, like word distribution, readability and formatting scores, the machine-learning system was often able to successfully predict the peer review outcome." For the WCNC dataset, the dense neural network achieved 74.01% accuracy and 72.30% F1-score compared to a random baseline of ~50%. For OpenReview, "75% of the samples have an error of under 1.2 (over a total score of 10), and the median error is only 0.79," demonstrating a strong correlation between superficial metrics and peer review decisions.
Reports effect sizes.
Research paradigm
Empirical, mixed-methods (quantitative machine learning + qualitative analysis)
Author conclusions
The authors conclude: "In this paper, we have reported an experiment involving three peer-reviewed conference proceedings, training a machine-learning system to infer a set of rules able to match the peer review outcome, ultimately providing an acceptance probability for other manuscripts. We focused on a rather superficial set of features of the submitted manuscripts, like word distribution, readability scores and document format. Nevertheless, the machine-learning system was often able to successfully predict the peer review outcome: we found a strong correlation between such superficial features and the outcome of the review process as a whole." They further state that "our findings point in the direction of a new type of analysis of typical human process, conducted with the help of machine-learning systems, one which is cognisant of ethical dimensions of the work as well as technical capabilities."
Risk of bias
Training data bias: Historical peer review decisions may encode existing sociocultural biases in the academic system; Algorithmic bias: AI systems trained on past data may propagate gender, language, and institutional affiliation biases; Geographic bias: Papers from low-income countries or underrepresented regions may be disadvantaged; Overfitting: Authors observed overfitting issues, such as specific first names being regarded as positive signals; Feature bias: Readability and formatting metrics may disproportionately disadvantage non-native English speakers; Design path bias: Models reflect designer values and goals frozen into code; Training bias: Models trained on past peer review decisions may replicate existing biases in historical review data; Geographic/institutional bias: Papers from under-represented regions or institutions may be disadvantaged; Language bias: Non-native English speakers may be penalized on readability metrics; Field-specific bias: Overfitting to conference-specific terminology (e.g., 'decoding' in machine learning conferences); First-impression bias: Superficial document features (formatting, readability) may unduly influence predictions; Selection bias in data collection: Data obtained from two sources (WCNC and OpenReview) with different review processes; Discrepancy between editorial decision and reviewer scores: Authors note "in a small number of cases, there is a rather strong discrepancy between the editorial decision and the average score of the reviewers"; Overfitting: Early stages showed overfitting with specific reviewer names treated as positive signals; Selection bias: Datasets limited to conference proceedings, primarily WCNC (wireless communications) and ICLR (machine learning); results may not generalize to journal submissions or other disciplines; Training data bias: Machine learning models trained on historical review decisions may perpetuate existing biases in the peer review process; Confounding: Papers with better formatting/readability may also have better scientific content; superficial features may be proxies for quality rather than causal drivers; Temporal bias: Models trained on past reviews may not reflect evolving submission quality from historically underrepresented regions; Geographic bias: Authors note USA dominates peer review (32.9% of reviews vs 25.4% of output); model may penalize papers from low-income countries; Overfitting risks: Specific keywords and author names initially flagged as predictive; document-dependent non-linear effects difficult to isolate
Limitations
- The approach does not cover all aspects of the peer review process, nor does it attempt to replace human reviewers with AI
- "The approach we take does not cover all the aspects of the peer review process, nor does it attempt to replace human reviewers with AI
- However, it suggests that there are some components of the quality assessment and peer review process, which could reasonably to assisted or replaced by AI-assisted tools." The authors note that "the possible model of quality assessment we are exploring is then a semi-automated one, where AI informs decision-making, rather than alone determining outcomes
- It is acknowledged the extent to which this is possible will vary depending on a number of factors, not least the nature of the research output itself and its approach to presenting research results." Additionally, the study does not cover complex intellectual tasks requiring domain expertise, and disciplinary variation may limit generalizability.
Open questions raised
- Feedback loop: Controlled experiments with academic reviewers to understand biases introduced by AI signals
- Review process: Incorporating full text of reviews and rebuttals (not just scores) to train AI tools
- Perception: Extended work on first-impression bias with more complex typographic layout indicators
- Disciplinary variation: How AI tools apply across different disciplinary norms
- Grant applications: Application of methods to grant review processes, which have different structure than research manuscripts
- Feedback loop: "We are interested in exploring the behaviour of reviewers when using these AI-powered support tools. We intend in future to carry out controlled experiments with academic reviewers, to understand the biases introduced by the AI signals on the reviewers."
Explore related topics
Related papers
- Comparing scientific abstracts generated by ChatGPT to real abstracts with detectors and blinded human reviewersCatherine A. Gao · 2023 · 657 citations
- Fabrication and errors in the bibliographic citations generated by ChatGPTWilliam H. Walters · 2023 · 352 citations
- Generative AI tools and assessment: Guidelines of the world's top-ranking universitiesBenjamin Luke Moorhouse · 2023 · 343 citations
- Human-AI collaboration patterns in AI-assisted academic writingAndy Nguyen · 2024 · 301 citations
- Hallucination Rates and Reference Accuracy of ChatGPT and Bard for Systematic Reviews: Comparative AnalysisMikaël Chelli · 2024 · 295 citations
- AI literacy and its implications for prompt engineering strategiesNils Knoth · 2024 · 277 citations