12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Human-in-the-Loop AI Reviewing: Feasibility, Opportunities, and Risks

Iddo Drori, Dov Te’eni · Journal of the Association for Information Systems · 2024

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
0/4
Quality (LMQS)
E
Evidence
38
Citations
3.93
FWCI
Top 10%
Impact

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.17705/1jais.00867

Methodology & findings

Study design

Experimental case study with comparative analysis.

Primary method

Comparative evaluation between LLM-generated reviews and human reviews; no specific statistical software mentioned in abstract.

Main result

The study found that "current AI-augmented reviewing is sufficiently accurate to alleviate the burden of reviewing but not completely and not for all cases." The authors demonstrated feasibility by evaluating and comparing LLM reviews with human reviews, and identified key opportunities and risks including bias, value misalignment, and misuse in AI-augmented academic peer review.

Reports effect sizes.

Research paradigm

Pragmatist/Mixed-methods (combines empirical experimentation with interpretive analysis)

Author conclusions

The authors conclude that "we explore the feasibility, opportunities, and risks of using large language models (LLMs) for reviewing academic submissions, while keeping the human in the loop" and that they "conclude with recommendations for managing these risks." They demonstrate that AI-augmented reviewing can alleviate reviewer burden while requiring human oversight to mitigate identified risks.

Risk of bias

Potential selection bias in choosing which submissions to review; Risk of algorithmic bias in GPT-4 reviews; Value misalignment between LLM outputs and reviewer standards; Representativeness of sample LLM reviews for broader applicability; Potential for LLM bias in academic reviewing; Value misalignment between AI systems and academic standards; Limited generalizability of GPT-4 performance across different submission types and domains; Bias in AI reviews (explicitly identified as a risk); Value misalignment between AI systems and human reviewers; Potential misuse of AI-augmented reviewing systems

Open questions raised

  • The authors present 'open questions' regarding opportunities of AI-augmented reviewing and recommend further work on risk management strategies for bias, value misalignment, and misuse in AI-assisted academic reviewing.
  • The authors present 'open questions' regarding AI-augmented reviewing and identify future directions related to managing risks of bias, value misalignment, and potential misuse of AI in academic peer review processes.
  • The authors present "open questions" regarding opportunities of AI-augmented reviewing and identify gaps in understanding how to manage risks of bias, value misalignment, and misuse in academic peer review contexts.
Data: not_statedCode: not_statedExtracted from: pdfAgreement 63%

Explore related topics

Related papers