12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

FactReview: Evidence-Grounded Peer Review with Execution-Based Claim Verification

Hang Xu, Ling Yue, Chaoqian Ouyang, Yuchen Liu, Libin Zheng, Shaowu Pan et al. · ArXiv.org · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
2/4
Quality (LMQS)
D
Evidence
0
Citations
0.00
FWCI

Methodology & findings

Study design

System design and comparative evaluation.

Primary method

Design science approach with iterative development and evaluation. System designed through multi-stage pipeline approach combining multiple evidence sources.

Main result

FactReview achieves the highest overall review-quality score of 4.86 out of 5 under an evidence-aware rubric, compared with 4.17 for DeepReview-v2 and 3.33 for matched OpenReview comments. The system "covers 84% of claims" across 35 ML papers and 463 benchmark claims. Crucially, "removing execution evidence changes 17% of claim statuses, more than any other single evidence source." In reviewer-assistance studies, "FactReview reduces mean review time by 58% while raising benchmark claim coverage from 87% to 99%."

Research paradigm

Design science / artifact development

Author conclusions

"FactReview treats automated peer review as evidence-grounded claim verification by coupling claim extraction, literature grounding, and execution-based verification. Across 35 ML papers and 463 benchmark claims, it covers 84% of benchmark claims and obtains the strongest evidence-aware review-quality score among the compared LLM systems and matched OpenReview comments." The authors argue that "LLM reviewers should audit empirical claims, not make accept-reject decisions" and that "FactReview is best understood as an audit layer for reviewers. It does not replace expert judgment, but makes the factual basis of that judgment easier to inspect."

Risk of bias

Limited sample size of 35 papers and small reviewer pool in assistance study may not generalize; Benchmark annotation bias: claims selected and labeled by annotators may not be representative of all review-relevant claims; Model selection bias: results specific to six general-purpose LLM backends tested; Venue/subfield bias: evaluation focused on empirical ML papers, limited to papers with released code; Execution environment bias: repair policy and resource budgets may differentially affect papers with different computational requirements; Limited to 35 papers and single detailed case study (limited generalization); Evidence-aware rubric reflects authors' judgment rather than venue-specific criteria; Small reviewer pool in assistance study (5 reviewers), may not generalize; Bounded repair policy may systematically underestimate reproducibility

Limitations

  • "Although our evaluation spans 35 ML papers and 463 benchmark claims, broader coverage of subfields, venues, and repository styles would be needed to fully characterize FactReview's behavior across diverse settings, and the qualitative depth analysis still relies on a single detailed case study
  • Second, execution-based verification is bounded by wall-clock time, compute budget, and a conservative repair policy: for experiments that require very long training, proprietary data, or large-scale infrastructure, FactReview may return Inconclusive even when the underlying claim is reproducible in principle
  • Third, our analysis covers six current general-purpose LLM backends
  • results may differ for models specialized in science or code, and absolute numbers will shift as models evolve
  • Fourth, the reviewer-assistance study uses a small reviewer pool on a fixed paper set, so the observed time and coverage gains may not generalize to larger reviewer populations."

Open questions raised

  • Future work should broaden reviewer and paper coverage, strengthen environment recovery and result alignment, and extend the evidence taxonomy beyond empirical ML papers. The paper notes that many errors come from execution-unavailable evidence, and remaining execution blockers concentrate in environment setup, runtime, metric availability, and claim alignment.
  • "Future work should broaden reviewer and paper coverage, strengthen environment recovery and result alignment, and extend the evidence taxonomy beyond empirical ML papers."
  • Future work should "broaden reviewer and paper coverage, strengthen environment recovery and result alignment, and extend the evidence taxonomy beyond empirical ML papers." The system could benefit from models "specialized in science or code" rather than general-purpose LLMs.
Data: Benchmark of 35 ML papers with 463 expert-annotated claims (availability not explicitly stated in text); 35 ML papers benchmark with 463 annotated claims (referenced in paper but no explicit URL provided); OpenReview comments from 24 papers (public OpenReview platform)Code: https://github.com/DEFENSE-SEU/FactReviewExtracted from: pdfAgreement 65%

Explore related topics

Related papers