12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Artificial intelligence to support publishing and peer review: A summary and review

Kayvan Kousha, Mike Thelwall · Learned Publishing · 2023

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
3/4
Quality (LMQS)
I
Evidence
137
Citations
58.57
FWCI
Top 10%
Impact

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1002/leap.1570

Methodology & findings

Study design

This is a narrative literature review and software summary.

Primary method

The review synthesizes multiple statistical approaches from cited studies including: machine learning algorithms (XGBoost, Random Forest, neural networks, deep learning), topic modeling, sentiment analysis, correlation analysis (Spearman's rho, Pearson correlations), and text similarity scoring. No original statistical analyses were performed in this review.

Main result

The study found that "automation is useful for helping to find reviewers and there is good evidence that it can sometimes help with initial quality control of submitted manuscripts." However, "The value of artificial intelligence (AI) to support reviewing has not been clearly demonstrated yet." Additionally, a study using the XGBoost algorithm achieved "an accuracy of 84% for academic journal recommendations" based on 20,250 articles, and another study using deep learning reported "an accuracy of 87%" for journal recommendations.

Reports effect sizes and confidence intervals.

Research paradigm

Interpretivist/critical review of existing literature and software systems

Author conclusions

The authors conclude: "The evidence above conclusively shows that AI is useful for helping to find reviewers and the spread of this technology to other contexts where it is not yet used, such as REF reviewer assignments, is recommended. There is also evidence that AI can sometimes support the initial quality control of submitted manuscripts. Although plagiarism detection is the obvious example, and is presumably widely used by publishers, statistical checking also seems useful and extending the capability of such software would be valuable. In contrast, there is insufficient evidence yet to use AI to support reviewing and it should not be used to replace human reviewers."

Risk of bias

Publication bias in reviewed literature (only published studies included); Selection bias in software selection for review (emphasis on documented systems); Geographic imbalance in reviewer populations affecting algorithm fairness; Gender imbalance in reviewer databases; Potential algorithmic bias in AI systems against certain writing styles or formatting; Publication bias: Only reviewed published academic evaluations; many tools lack published efficacy data; Selective coverage: Focused on tools with academic research; excluded many commercial systems without published evaluations; Literature scope: Limited to peer review and publishing AI applications; not comprehensive across all domains

Limitations

  • "The availability of open peer review reports varies substantially between journals and fields (Ortega, 2019), and between reviewer countries (Severin et al., 2021), further undermining its value as an input for AI." Additionally, the authors note that "peer review text and scores are currently too sparsely available to support post-publication research assessment" and that "Much other software supporting publishing and editorial work exists and is being used, but without published academic evaluations of its efficacy."

Open questions raised

  • AI capability for reviewing has not been clearly demonstrated and further testing is important
  • Limited published academic evaluations of many commercial software systems
  • Need for robust accuracy measures for automated reviewer assignment systems
  • Insufficient availability of peer review text and scores for post-publication research assessment
  • Lack of AI analyses of post-publication peer review to date
  • Need to test AI systems for desk rejection of obviously poor papers
Data: PeerRead dataset of scientific peer reviews (mentioned for review decision prediction); F1000Research reviewer decisions (used for PeerJudge validation); MDPI journals open peer review data (45,385 reviews across 288 journals)Code: ReviewAdvisor - https://github.com/neulab/ReviewAdvisor; PeerJudge - http://sentistrength.wlv.ac.uk/PeerJudge.htmlExtracted from: pdfAgreement 66%

Explore related topics

Related papers