PRISM: A Multi-Dimensional Benchmark for Evaluating LLM Peer Reviewers
Ngoc Phan Phuoc Loc, Toan Huynh La Viet, Thanh Tran Khanh, Duy Anh Nguyễn, Tuan Anh Nguyen Pham, Thành Nguyen et al. · ArXiv.org · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
Methodology & findings
Study design
Multi-dimensional benchmarking study using four dedicated evaluation pipelines: (1) Depth of Analysis via argument mining with grounding-level classification; (2) Novelty Assessment using retrieval-augmented verification against Semantic Scholar; (3) Flaw Identification & Major Issues Prioritization via consensus-based verification and positional ranking; (4) Multi-dimensional Constructiveness via atomic review comment scoring across five dimensions (actionability, specificity, justification, solution orientation, tone).
Main result
The study found that "no single system consistently matches the balanced performance of the human baseline across all dimensions at once." LLMs demonstrated distinct specialization profiles: "CycleReviewer and DeepReview match human analytical depth; TreeReview falls into a surface-level trap, overindexing on presentation anomalies." On novelty assessment, "SEA-E outperforms human reviewers on grounded novelty verification; other systems exhibit measurable novelty hallucination." For flaw identification, "Reviewer2 leads in flaw recall as a high-sensitivity scanner; LLMs broadly achieve near-perfect critical issue prioritization, demonstrating a cognitive alignment comparable to human reviewers." Regarding constructiveness, "DeepReview produces the most actionable feedback, though a constructiveness gap relative to human reviewers persists across all systems."
Research paradigm
Empirical-Computational (benchmarking framework with systematic evaluation)
Author conclusions
"PRISM demonstrates that LLM peer reviewers are specialized tools rather than general-purpose replacements for human expertise. Each system excels in a specific niche but exhibits distinct blind spots across other dimensions." The authors recommend "a targeted ensemble deployment rather than a standalone approach: use Reviewer2 for exhaustive flaw scanning (highest diagnostic recall); use DeepReview for constructive feedback drafting (highest actionability and solution density); use SEA for novelty-grounding checks (highest literature support rate). Ultimately, these systems are most effective as specialist co-pilots within a human-assisted pipeline rather than autonomous reviewers."
Risk of bias
Judge-specific bias: Reliance on Gemini 2.5 Flash Lite as sole evaluation engine for all metric extraction and scoring tasks; Domain specificity bias: Benchmark corpus limited to ML/AI venues (ICLR, ICML, NeurIPS) only; Corpus composition bias: Stratified by decision category but may not represent full distribution of scientific papers; Retrieval bias: Semantic Scholar API used for novelty verification may have coverage gaps or ranking biases; Ground truth bias: Consensus mechanism for flaw identification merges human and LLM reviewers, potentially conflating their judgments; Selection bias: Figures and visual elements excluded from evaluation as LLM reviewers lack reliable multimodal support; Judge-specific bias: reliance on single LLM (Gemini 2.5 Flash Lite) for evaluation; Domain-specific bias: benchmark limited to ML/AI venues (ICLR, ICML, NeurIPS); Venue-specific score distribution: sampling preserves original venue acceptance distributions which may not generalize; Hallucination risk in LLM-as-Judge for flaw verification and novelty claim scoring; Potential position bias in prioritization scoring based on review text ordering; Judge-specific biases from reliance on Gemini 2.5 Flash Lite as sole evaluation engine; Venue-specific bias: evaluation limited to machine learning conferences only; Domain generalization risk: unclear if PRISM metrics transfer to other scientific fields; Potential introduction of position bias in flaw prioritization scoring; Hallucination effects in flaw extraction, particularly for minor flaws
Limitations
- The authors state: "Our primary evaluation pipeline relies on Gemini 2.5 Flash Lite as the core judge model
- While we conducted preliminary robustness checks using an alternative model (Xiaomi MiMo V2.5 Pro) on a subset of the data to verify metric stability, a comprehensive multi-judge study across diverse LLM families remains necessary to fully eliminate judge-specific biases
- Furthermore, the benchmark corpus covers ML/AI venues only, and PRISM may require recalibration for other scientific domains."
Open questions raised
- Cross-domain generalization: PRISM requires recalibration for clinical medicine, social sciences, and pure mathematics beyond ML/AI venues
- Judge robustness: Systematic study of inter-judge agreement across diverse LLM judge families and human raters needed
- Human validation: Correlation of PRISM scores with post-review author satisfaction or acceptance decision outcomes to confirm metrics capture meaningful review quality
- Cross-domain generalization: recalibrating PRISM for clinical medicine, social sciences, and pure mathematics
- Judge robustness: systematic study of inter-judge agreement across LLM judge families and human raters
- Human validation: correlating PRISM scores with post-review author satisfaction or acceptance decision outcomes to confirm metrics capture meaningful review quality
Explore related topics
Related papers
- Which academic search systems are suitable for systematic reviews or meta‐analyses? Evaluating retrieval qualities of Google Scholar, PubMed, and 26 other resourcesMichael Gusenbauer · 2019 · 2,116 citations
- Generative AI tools and assessment: Guidelines of the world's top-ranking universitiesBenjamin Luke Moorhouse · 2023 · 343 citations
- AI-assisted peer reviewAlessandro Checco · 2021 · 261 citations
- Fighting reviewer fatigue or amplifying bias? Considerations and recommendations for use of ChatGPT and other large language models in scholarly peer reviewMohammad Hosseini · 2023 · 209 citations
- Artificial intelligence to support publishing and peer review: A summary and reviewKayvan Kousha · 2023 · 137 citations
- Artificial intelligence adoption in the physical sciences, natural sciences, life sciences, social sciences and the arts and humanities: A bibliometric analysis of research publications from 1960-2021Stefan Hajkowicz · 2023 · 119 citations