12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

PRISM: A Multi-Dimensional Benchmark for Evaluating LLM Peer Reviewers

Ngoc Phan Phuoc Loc, Toan Huynh La Viet, Thanh Tran Khanh, Duy Anh Nguyễn, Tuan Anh Nguyen Pham, Thành Nguyen et al. · ArXiv.org · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
2/4
Quality (LMQS)
C
Evidence
0
Citations
0.00
FWCI

Methodology & findings

Study design

Multi-dimensional benchmarking study using four dedicated evaluation pipelines: (1) Depth of Analysis via argument mining with grounding-level classification; (2) Novelty Assessment using retrieval-augmented verification against Semantic Scholar; (3) Flaw Identification & Major Issues Prioritization via consensus-based verification and positional ranking; (4) Multi-dimensional Constructiveness via atomic review comment scoring across five dimensions (actionability, specificity, justification, solution orientation, tone).

Main result

The study found that "no single system consistently matches the balanced performance of the human baseline across all dimensions at once." LLMs demonstrated distinct specialization profiles: "CycleReviewer and DeepReview match human analytical depth; TreeReview falls into a surface-level trap, overindexing on presentation anomalies." On novelty assessment, "SEA-E outperforms human reviewers on grounded novelty verification; other systems exhibit measurable novelty hallucination." For flaw identification, "Reviewer2 leads in flaw recall as a high-sensitivity scanner; LLMs broadly achieve near-perfect critical issue prioritization, demonstrating a cognitive alignment comparable to human reviewers." Regarding constructiveness, "DeepReview produces the most actionable feedback, though a constructiveness gap relative to human reviewers persists across all systems."

Research paradigm

Empirical-Computational (benchmarking framework with systematic evaluation)

Author conclusions

"PRISM demonstrates that LLM peer reviewers are specialized tools rather than general-purpose replacements for human expertise. Each system excels in a specific niche but exhibits distinct blind spots across other dimensions." The authors recommend "a targeted ensemble deployment rather than a standalone approach: use Reviewer2 for exhaustive flaw scanning (highest diagnostic recall); use DeepReview for constructive feedback drafting (highest actionability and solution density); use SEA for novelty-grounding checks (highest literature support rate). Ultimately, these systems are most effective as specialist co-pilots within a human-assisted pipeline rather than autonomous reviewers."

Risk of bias

Judge-specific bias: Reliance on Gemini 2.5 Flash Lite as sole evaluation engine for all metric extraction and scoring tasks; Domain specificity bias: Benchmark corpus limited to ML/AI venues (ICLR, ICML, NeurIPS) only; Corpus composition bias: Stratified by decision category but may not represent full distribution of scientific papers; Retrieval bias: Semantic Scholar API used for novelty verification may have coverage gaps or ranking biases; Ground truth bias: Consensus mechanism for flaw identification merges human and LLM reviewers, potentially conflating their judgments; Selection bias: Figures and visual elements excluded from evaluation as LLM reviewers lack reliable multimodal support; Judge-specific bias: reliance on single LLM (Gemini 2.5 Flash Lite) for evaluation; Domain-specific bias: benchmark limited to ML/AI venues (ICLR, ICML, NeurIPS); Venue-specific score distribution: sampling preserves original venue acceptance distributions which may not generalize; Hallucination risk in LLM-as-Judge for flaw verification and novelty claim scoring; Potential position bias in prioritization scoring based on review text ordering; Judge-specific biases from reliance on Gemini 2.5 Flash Lite as sole evaluation engine; Venue-specific bias: evaluation limited to machine learning conferences only; Domain generalization risk: unclear if PRISM metrics transfer to other scientific fields; Potential introduction of position bias in flaw prioritization scoring; Hallucination effects in flaw extraction, particularly for minor flaws

Limitations

  • The authors state: "Our primary evaluation pipeline relies on Gemini 2.5 Flash Lite as the core judge model
  • While we conducted preliminary robustness checks using an alternative model (Xiaomi MiMo V2.5 Pro) on a subset of the data to verify metric stability, a comprehensive multi-judge study across diverse LLM families remains necessary to fully eliminate judge-specific biases
  • Furthermore, the benchmark corpus covers ML/AI venues only, and PRISM may require recalibration for other scientific domains."

Open questions raised

  • Cross-domain generalization: PRISM requires recalibration for clinical medicine, social sciences, and pure mathematics beyond ML/AI venues
  • Judge robustness: Systematic study of inter-judge agreement across diverse LLM judge families and human raters needed
  • Human validation: Correlation of PRISM scores with post-review author satisfaction or acceptance decision outcomes to confirm metrics capture meaningful review quality
  • Cross-domain generalization: recalibrating PRISM for clinical medicine, social sciences, and pure mathematics
  • Judge robustness: systematic study of inter-judge agreement across LLM judge families and human raters
  • Human validation: correlating PRISM scores with post-review author satisfaction or acceptance decision outcomes to confirm metrics capture meaningful review quality
Data: Benchmark corpus of 1,000 papers and reviews from ICLR (2024-2026), ICML (2025), and NeurIPS (2025) - available at https://prism-benchmark.github.io/; 1,000 papers from ICLR 2024-2026, ICML 2025, NeurIPS 2025 with stratified sampling across decision categories (Oral, Spotlight, Poster, Reject); 5,000 automated reviews generated from five LLM reviewer systems; Human reviewer baseline reviews from official conference review pools; 1,000 papers corpus from ICLR 2024-2026, ICML 2025, NeurIPS 2025 with stratification by decision category; Demo and key results available at https://prism-benchmark.github.io/Code: Demo and key results: https://prism-benchmark.github.io/ (specific GitHub/code repository URL not explicitly provided in paper); https://prism-benchmark.github.io/ (demo and key results); https://prism-benchmark.github.io/ (demo and results)Extracted from: pdfAgreement 48%

Explore related topics

Related papers