12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

What Makes a Good AI Review? Concern-Level Diagnostics for AI Peer Review

Ming Jin · ArXiv.org · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
1/4
Quality (LMQS)
D
Evidence
0
Citations
0.00
FWCI

Methodology & findings

Study design

Framework demonstration via concern alignment evaluation ladder applied to 48 papers from ICLR 2026, NeurIPS 2025, and ICML 2025 in AI safety/alignment domain (24 accepted, 24 rejected).

Primary method

Design science research with iterative refinement. The paper develops a concern-alignment diagnostic framework organized as an evaluation ladder with five levels (L0-L4).

Main result

The study reveals that "identical verdict accuracy can conceal reject-heavy and low-recall profiles; that most single-agent systems mark 25-55% of concerns on accepted papers as decisive; that changing the model while holding prompts fixed shifts reviewer behavior in measurable, non-uniform ways; and that some systems attend as readily to resolved concerns as to decisive blockers." Detection alone does not determine review quality; calibration is often the binding constraint, with systems frequently over-escalating concern severity despite correctly identifying underlying issues.

Research paradigm

Design science / evaluation framework development

Author conclusions

"Binary accuracy cannot tell us how a system fails or what to fix. Concern alignment does so by evaluating the unit of review that authors and readers actually act on." The authors emphasize that "each level of the ladder exposed a different failure mode: verdict stratification revealed reject-heavy behavior; decision-aware metrics quantified severity miscalibration; rebuttal-aware recall exposed inverted attention; and top-K analysis separated calibrated prioritization from concern dilution." Most importantly, "The clearest lesson is that finding issues is not enough. The systems we examined often detect real concerns yet still misjudge their decision weight, especially on accepted papers."

Risk of bias

Selection bias: Papers intentionally selected from AI safety/alignment domain with high-quality reviews; Measurement bias: LLM-assisted match graph construction introduces potential circularity despite cross-model validation; Severity inference bias: For Systems L, A, and O, severity is inferred by extraction pipeline rather than emitted natively by systems; Verdict inference bias: Four of six configurations lack explicit accept/reject decisions; verdicts inferred from tone using default-REJECT rule; Anchor bias: Framework uses noisy post-deliberation AC decisions as operational anchor for calibrating concern severity; Sample composition bias: 22 of 24 rejected papers intentionally selected as 'hard negatives' with substantive reasons for rejection; Selection bias: papers filtered to AI safety/alignment domain controlling for reviewing norms; Measurement bias: severity assignment inferred by extraction pipeline for 3 of 4 systems rather than natively emitted; Circularity risk: match graphs constructed using LLMs to evaluate LLM-generated reviews; Anchor bias: AC decisions used as operational ground truth despite acknowledged noisiness; Verdict inference bias: four of six system configurations lack explicit accept/reject decisions; verdicts inferred from review tone using default-REJECT rule; LLM-assisted match graph construction creates circularity risk (mitigated via cross-model validation with different model family, human adjudication, and calibration exemplars); Verdict inference methodology sensitive to interpretation method (46-96 percentage point swing across methods per Appendix U audit); AC decisions used as operational anchor are themselves noisy; accepted papers may contain genuine blockers, rejected papers may lack clear justification; Sample limited to AI safety/alignment domain; selection prioritized review quality (≥3 substantive reviews, unambiguous AC decision, ≥2 extractable concerns), potentially excluding lower-quality reviews; Severity assignment by extraction pipeline for Systems L, A, O (not natively emitted) means Level 3-4 metrics partly evaluate inference step; System M multi-agent output structure makes verdict inference unreliable regardless of method (all 48 reviews flagged as containing coordination artifacts); Evaluator independence limited: match graph construction uses LLM; semantic verification also LLM-based but different model family

Limitations

  • "Our pilot covers 48 papers in one domain and four systems
  • That is enough to expose failure modes hidden by verdict accuracy, but not enough for population-level claims or fine-grained rankings." Additionally, "Match graphs are constructed with LLM assistance, so circularity remains a methodological risk," and "Because verdict accuracy numbers for Systems L, O, and M reflect this inference procedure rather than native system decisions, we conducted a verdict inference audit" which showed "verdict-level findings are sensitive to the inference method, while the paper's concern-level diagnostics..
  • are unchanged by how accept/reject is inferred." The authors note that "System M was run only on GPT-4o, so its profile remains a combined model-plus-method effect."

Open questions raised

  • Population-level claims require larger sample beyond 48 papers in single domain
  • Fine-grained system rankings require additional evaluation scope
  • Cross-domain assessment needed beyond AI safety/alignment topic area
  • Relationship between verdict readability and calibration requires further investigation
  • Transfer validity of phantom-quality audit across different system types
  • Formal theoretical framework for the evaluation ladder is lacking (currently heuristic)
Data: OpenReview dataset for papers from ICLR 2026, NeurIPS 2025, ICML 2025 in AI safety/alignment domain (48 papers total; identities to be released with supplementary data); Paper sourcing from OpenReview across ICLR 2026, NeurIPS 2025, and ICML 2025. The paper states: "Papers were sourced from OpenReview and filtered to the AI safety/alignment domain to control for topic-specific reviewing norms. Selection prioritized review quality: each paper has ≥3 substantive reviews, an unambiguous AC decision with articulated reasoning, and ≥2 extractable technical concerns." Data curation details provided in Appendix P.; Evaluation set of 48 papers from ICLR 2026, NeurIPS 2025, ICML 2025 available for supplementary release with identities. Official concern sheets and match graphs mentioned as auditable artifacts but specific public dataset URL not provided in text.Code: System L: https://github.com/allenai/marg-reviewer; System A: https://github.com/SakanaAI/AI-Scientist; System O: https://github.com/ChicagoHAI/OpenAIReview; System L (Liang et al., 2024b): MARG repository; System A (Lu et al., 2024): AI Scientist repository; System O (ChicagoHAI, 2026): openaireview pip package (v0.2.7); System M (D'Arcy et al., 2024): source repository referenced but implementation details in Appendix D; System A (AI Scientist): https://github.com/SakanaAI/AI-Scientist; System O (ChicagoHAI OpenAIReview): https://github.com/ChicagoHAI/OpenAIReviewExtracted from: pdfAgreement 42%

Explore related topics

Related papers