FactReview: Evidence-Grounded Peer Review with Execution-Based Claim Verification
Hang Xu, Ling Yue, Chaoqian Ouyang, Yuchen Liu, Libin Zheng, Shaowu Pan et al. · ArXiv.org · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
Methodology & findings
Study design
System design and comparative evaluation.
Primary method
Design science approach with iterative development and evaluation. System designed through multi-stage pipeline approach combining multiple evidence sources.
Main result
FactReview achieves the highest overall review-quality score of 4.86 out of 5 under an evidence-aware rubric, compared with 4.17 for DeepReview-v2 and 3.33 for matched OpenReview comments. The system "covers 84% of claims" across 35 ML papers and 463 benchmark claims. Crucially, "removing execution evidence changes 17% of claim statuses, more than any other single evidence source." In reviewer-assistance studies, "FactReview reduces mean review time by 58% while raising benchmark claim coverage from 87% to 99%."
Research paradigm
Design science / artifact development
Author conclusions
"FactReview treats automated peer review as evidence-grounded claim verification by coupling claim extraction, literature grounding, and execution-based verification. Across 35 ML papers and 463 benchmark claims, it covers 84% of benchmark claims and obtains the strongest evidence-aware review-quality score among the compared LLM systems and matched OpenReview comments." The authors argue that "LLM reviewers should audit empirical claims, not make accept-reject decisions" and that "FactReview is best understood as an audit layer for reviewers. It does not replace expert judgment, but makes the factual basis of that judgment easier to inspect."
Risk of bias
Limited sample size of 35 papers and small reviewer pool in assistance study may not generalize; Benchmark annotation bias: claims selected and labeled by annotators may not be representative of all review-relevant claims; Model selection bias: results specific to six general-purpose LLM backends tested; Venue/subfield bias: evaluation focused on empirical ML papers, limited to papers with released code; Execution environment bias: repair policy and resource budgets may differentially affect papers with different computational requirements; Limited to 35 papers and single detailed case study (limited generalization); Evidence-aware rubric reflects authors' judgment rather than venue-specific criteria; Small reviewer pool in assistance study (5 reviewers), may not generalize; Bounded repair policy may systematically underestimate reproducibility
Limitations
- "Although our evaluation spans 35 ML papers and 463 benchmark claims, broader coverage of subfields, venues, and repository styles would be needed to fully characterize FactReview's behavior across diverse settings, and the qualitative depth analysis still relies on a single detailed case study
- Second, execution-based verification is bounded by wall-clock time, compute budget, and a conservative repair policy: for experiments that require very long training, proprietary data, or large-scale infrastructure, FactReview may return Inconclusive even when the underlying claim is reproducible in principle
- Third, our analysis covers six current general-purpose LLM backends
- results may differ for models specialized in science or code, and absolute numbers will shift as models evolve
- Fourth, the reviewer-assistance study uses a small reviewer pool on a fixed paper set, so the observed time and coverage gains may not generalize to larger reviewer populations."
Open questions raised
- Future work should broaden reviewer and paper coverage, strengthen environment recovery and result alignment, and extend the evidence taxonomy beyond empirical ML papers. The paper notes that many errors come from execution-unavailable evidence, and remaining execution blockers concentrate in environment setup, runtime, metric availability, and claim alignment.
- "Future work should broaden reviewer and paper coverage, strengthen environment recovery and result alignment, and extend the evidence taxonomy beyond empirical ML papers."
- Future work should "broaden reviewer and paper coverage, strengthen environment recovery and result alignment, and extend the evidence taxonomy beyond empirical ML papers." The system could benefit from models "specialized in science or code" rather than general-purpose LLMs.
Explore related topics
Related papers
- Generative AI tools and assessment: Guidelines of the world's top-ranking universitiesBenjamin Luke Moorhouse · 2023 · 343 citations
- AI-assisted peer reviewAlessandro Checco · 2021 · 261 citations
- Fighting reviewer fatigue or amplifying bias? Considerations and recommendations for use of ChatGPT and other large language models in scholarly peer reviewMohammad Hosseini · 2023 · 209 citations
- Artificial intelligence to support publishing and peer review: A summary and reviewKayvan Kousha · 2023 · 137 citations
- Ethical Dilemmas in Using AI for Academic Writing and an Example Framework for Peer Review in Nephrology Academia: A Narrative ReviewJing Miao · 2023 · 93 citations
- Artificial Intelligence in Peer Review: Enhancing Efficiency While Preserving IntegrityBohdana Doskaliuk · 2025 · 59 citations