12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Benchmarking Agentic Review Systems

Dang Van Nguyen, Wanqing Hao, Yanai Elazar, Chenhao Tan · ArXiv.org · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
2/4
Quality (LMQS)
C
Evidence
0
Citations
0.00
FWCI

Methodology & findings

Study design

Multi-part benchmarking study combining three evaluation approaches: (1) correlation study on ICLR/NeurIPS papers (2017–2025) measuring whether AI-generated comment volume correlates with paper quality proxies (citations, awards, review scores); (2) controlled perturbation benchmark injecting four error categories into papers across eight arXiv subject classes and measuring error detection recall; (3) observational deployment study collecting user feedback on 1,360 reviews of 1,100 papers via a public web tool..

Main result

The study found that "the best is OpenAIReview + GPT5.5 at 83.0%" pairwise accuracy on ICLR/NeurIPS papers, and "The strongest configuration (OpenAIReview + GPT-5.5) catches 71.6% of injected errors," demonstrating that while AI review systems can track human quality judgments and catch important errors, "they still have room for improvement." Additionally, "the union of detections across six models reaches 83.3% recall, suggesting different models detect different errors and better harness design can potentially increase performance."

Research paradigm

Empirical-computational; pragmatic evaluation of AI systems in practice

Author conclusions

The authors conclude that "while AI reviews still have room for improvement, they can already track human quality judgments well, catch important errors, and earn positive feedback from real users." They state that "today's AI reviewer systems pick up on signals of paper quality at 66−83% accuracy without being explicitly trained to, with a stronger signal in stronger models." Furthermore, they note that "AI review systems are positioned to serve peer review as a complement to human reviewers, handling the exhaustive enumeration of concrete issues while human reviewers focus on the higher-level judgments where they remain the dominant source of feedback."

Risk of bias

Selection bias in quality-proxy construction: papers deliberately selected from top/bottom groups rather than uniform samples; Citation bias: uses citations-per-year as quality signal which may favor certain research areas and topics; Generator LLM bias in perturbation selection: audit found 'only modest bias relative to random selection, mainly a preference for longer equations'; Manual validation ambiguity: 12.5% of perturbations rated as ambiguous in manual audit; User deployment bias: anonymous web users likely skewed toward researchers in computer science and AI; Selection bias: Quality proxy construction uses deliberately selected contrasts (top/bottom groups) rather than uniform samples, which may not represent full distribution; Model selection bias in perturbation generation: Generator LLM audit found "modest bias relative to random selection, mainly a preference for longer equations" (Appendix C.4); Sampling design: Frontier models only evaluated on 74-paper subset (24 papers for some analyses) rather than full dataset due to cost constraints; Observational deployment bias: Public web tool users "most likely researchers" and papers "span many fields" with uncontrolled selection mechanism; Proxy validity bias: Citations, awards, and review scores are acknowledged as noisy and may not reflect true paper quality; Error detection evaluation bias: Fuzzy substring matching and LLM judge criteria (threshold τ and cutoff values in Appendix C.6) introduce subjective classification; Quality proxies (citations, awards, review scores) are noisy and may not reflect true paper quality; Sample selection bias: ICLR/NeurIPS papers may not represent all research fields or quality distributions; Generator LLM bias in perturbation selection shows 'preference for longer equations' (Appendix C.4); User deployment bias: anonymous web users likely skew toward computer science and AI researchers; Manual audit of perturbations found 5 ambiguous cases (12.5%) and 2 invalid cases (5%)

Limitations

  • The authors acknowledge that "Pairwise accuracy on comment counts is admittedly a coarse metric, and may miss cases where a concise review points out a single fatal issue." Additionally, the quality proxies themselves are acknowledged as "noisy proxies of quality" and the authors emphasize they "select the top and bottom groups as a tractable approximation, not as ground-truth measures of paper quality." Regarding the perturbation benchmark, manual audit found "2 (5%) to be not true errors, and 5 (12.5%) to be ambiguous, with the latter two categories concentrated in the surface-numeric and empirical-claim subtypes." Cost limitations led to restricting frontier model evaluation: "Running the two frontier models (GPT-5.5 and Claude Opus 4.7) under the zero-shot and OpenAIReview methods on the 80-paper subset cost approximately $204 in API charges."

Open questions raised

  • Improving error detection recall beyond 71.6% for frontier models and 83.3% for model ensembles
  • Reducing false positives and minor nitpicks - identified as main user complaints in deployment
  • Improving precision of comment detection
  • Better harness design to achieve higher complementarity across models
  • Better handling of surface-level math errors (showing only modest gains from OpenAIReview design)
  • Extending evaluation beyond ICLR/NeurIPS to broader range of research areas
Data: SNOR dataset (Neumann, 2025) linking OpenReview submissions for ICLR (2017–2025) and NeurIPS (2021–2025) to Semantic Scholar; 74 arXiv papers from 8 subject classes used in perturbation benchmark; SNOR dataset: Links OpenReview submissions for ICLR (2017–2025) and NeurIPS (2021–2025) to Semantic Scholar with citation counts, decisions, and reviewer scores (Neumann, 2025); Perturbation benchmark: 74 papers from 8 arXiv subject classes (Computational Complexity, Machine Learning, Econometrics, Experimental High-Energy Physics, Mathematics, Atomic and Cluster Physics, Genomics, Applied Statistics) with injected errors and ground truth annotations; SNOR dataset (Neumann, 2025) linking OpenReview submissions to Semantic Scholar citation counts and decisionsCode: OpenAIReview: Open-source system (Chicago Human+AI Lab, 2025) - exact repository URL not provided in paper; 'coarse: Open-source multi-agent reviewer (Van Dijcke, 2025) - exact repository URL not provided; Reviewer3: Proprietary commercial system (Reviewer3, 2025) - code not available; OpenAIReview (open-source, Chicago Human+AI Lab, 2025)Extracted from: pdfAgreement 50%

Explore related topics

Related papers