12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Enhancing peer review efficiency: A mixed‐methods analysis of artificial intelligence‐assisted reviewer selection across academic disciplines

Shai Farber · Learned Publishing · 2024

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
3/4
Quality (LMQS)
E
Evidence
29
Citations
19.75
FWCI
Top 10%
Impact

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1002/leap.1638

Methodology & findings

Study design

Mixed-methods study combining quantitative measurement and qualitative analysis.

Sample

N = 20, 4 groups

Primary method

Paired t-tests conducted to compare time efficiency between AI and traditional methods. Descriptive statistics calculated for selection accuracy, quality, and editor satisfaction ratings. Thematic analysis using constant comparative method (Glaser & Strauss, 1967) for qualitative data with two independent coders and inter-rater reliability checks. Subgroup analyses performed to compare AI performance across different academic fields (STEM vs. Social Sciences & Humanities). Effect sizes (Cohen's d) calculated for disciplinary differences.

Main result

The study found that "the AI system achieved a 42% overlap with editors' selections and demonstrated a significant improvement in time efficiency, reducing selection time by 73%". Additionally, "Editors found that 37% of AI-suggested reviewers who were not part of their initial selection were indeed suitable". The system's performance varied significantly across disciplines: "The system's performance varied across disciplines, with higher accuracy in STEM fields (Cohen's d = 0.68)".

Reports effect sizes and confidence intervals.

Research paradigm

Pragmatist (mixed-methods combining quantitative and qualitative approaches)

Author conclusions

The authors conclude that "The AI system achieved a 42% overlap with editors' choices, illustrating both its promise and inherent limitation. Notably, the system's performance varied across disciplines, performing more robustly in STEM fields compared to the social sciences and humanities." They emphasize that "Perhaps the most notable finding is the 73% reduction in reviewer selection time, a substantial efficiency gain with far-reaching implications for reducing the burden on editors and accelerating the peer-review process." However, they caution that "these promising results must be tempered with an acknowledgment of the system's limitations. The occasional suggestion of overly prominent or, in rare cases, fictional reviewers highlight the ongoing need for human oversight and the importance of editors' expertise in the reviewer selection process."

Risk of bias

Selection bias: snowball sampling method for editor recruitment may not represent diverse journal populations; Recall bias: time measurements based on editor self-reporting; Single manuscript per editor: does not capture variability across different manuscript types; AI training data bias: potential geographic and institutional disparities, skew towards North American and European institutions; Algorithmic bias: AI models may perpetuate existing academic biases in training data; Verification bias: AI-generated reviewer affiliations not systematically verified; Snapshot limitation: single time point assessment (April 2024) may not reflect AI performance changes; Selection bias: Snowball sampling method with editors recruited through personal acquaintance may not represent diverse editorial perspectives; Algorithmic bias: AI training data may perpetuate geographic and institutional biases, with perceived skew towards researchers from North American and European institutions; Recall bias: Time efficiency data based on editor self-reporting; Single-instance evaluation: Each editor provided only one manuscript evaluation, not capturing full spectrum of selection challenges; Potential demographic bias in AI training data; Gender imbalance in publication datasets may be reflected in AI recommendations; Lack of systematic verification of reviewer affiliations provided by AI; Selection bias: Snowball sampling method may not be representative; Recall bias: Time efficiency data relied on self-reporting; Algorithmic bias: AI training data may perpetuate existing academic biases favoring North American and European institutions; Single-instance evaluation: Potential instance-specific biases from evaluating only one manuscript per editor; Geographic/institutional bias: Perceived skew towards researchers from well-known institutions (10% of editors noted this concern); Demographic representation bias in AI training data, particularly underrepresentation of non-English scholarship

Limitations

  • The authors acknowledge several key limitations: "While the study included 20 editors from various disciplines, this may not fully represent the global publishing landscape, constraining the generalizability of the findings"
  • "The study provides only a snapshot of AI performance, not assessing long-term impacts on the peer-review process or publication quality"
  • "The AI system's training data, while extensive, may not fully represent the diversity of global scholarship, especially from regions with less digitized or accessible research outputs
  • This could lead to systemic underrepresentation of scholars from certain geographical areas or institutions in reviewer recommendations." Additionally, "time efficiency data relied on editors' self-reporting, which may be subject to recall bias or imprecision"
  • The authors also note "the questionnaire used has not undergone formal validation, which is common in initial investigations of new phenomena".

Open questions raised

  • Limited large-scale empirical studies comparing AI-assisted reviewer selection with traditional methods across diverse disciplines
  • Long-term impacts of AI integration on peer-review quality and publication outcomes remain underexplored
  • Limited research on perceptions and attitudes of editors, reviewers, and authors towards AI-assisted peer review
  • Development of ethical frameworks and best practices for responsible AI use in peer review still in infancy
  • AI performance in interdisciplinary fields not fully explored
  • Need for inter-editor agreement rate studies to establish baseline human expert performance
Extracted from: pdfAgreement 59%

Explore related topics

Related papers