12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Review the Code, Not the Story: A Vision and Protocol for Code-First Peer Review

Jienan Chen · arXiv (Cornell University) · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
1/4
Quality (LMQS)
I
Evidence
0
Citations
0.00
FWCI

Methodology & findings

Study design

Vision and protocol proposal using worked examples, threat modeling, and design goals.

Main result

The paper proposes that "code-first peer review" can shift computational research evaluation from author-controlled manuscripts to "venue-controlled" evidence packages generated from executable artifacts. The authors argue that "Instead of reading an author-controlled narrative first and code second, reviewers can read a venue-generated evidence view that is grounded in executed artifacts." Current AI agents achieve only limited reproducibility—"the best tested agent achieved a 21.0% average replication score" (PaperBench) and "the best agent achieved 21% accuracy on the hardest task" (CORE-Bench)—demonstrating that AI should assist with evidence extraction rather than replace human judgment.

Research paradigm

Normative/prescriptive (proposal for infrastructure redesign)

Author conclusions

The authors conclude: "Computational peer review should review executable evidence before reviewing polished stories. We proposed code-first peer review, a venue-controlled protocol in which authors submit artifacts and minimal claim manifests, AI systems execute and audit those artifacts, and reviewers evaluate standardized evidence packages." They further state: "By shifting control from author-crafted narrative to venue-generated evidence views, code-first review can improve reproducibility, expose hidden implementation details, reduce reviewer burden, and make unsupported claims easier to detect. The appropriate response is not to avoid AI in review infrastructure, but to govern it: blind metadata, audit models, log provenance, preserve author appeal, and keep human reviewers responsible for scientific significance and final decisions."

Risk of bias

Model bias in AI-assisted review (acknowledged as central governance concern); Affiliation bias in LLM-mediated review (reported 38.4% to 40.0% acceptance rate variation by affiliation tier in JAMA study); Metadata bias (scores changing with author institution or name); Model-version drift (different versions producing different claim statuses); Over-standardization penalizing nontraditional contributions; AI model bias in generated review views; Affiliation bias in LLM-mediated review (reported 38.4% acceptance for top-tier vs 36.7% for low-tier affiliations); Metadata leakage through code repositories (usernames, commit history, file paths, dataset URLs); Model-version drift leading to inconsistent review standards; Hidden prompt injection in submitted artifacts to manipulate AI decisions; Metadata bias from author affiliations and identities (mitigated by blinding); Model-version drift across evaluation runs; Prompt-injection attacks via submitted text; Potential over-standardization disadvantaging non-traditional contributions; False accusation risks when AI flags defects incorrectly

Limitations

  • The authors explicitly state: "This paper is a vision and protocol proposal, not a completed deployed system
  • Several open questions remain
  • First, current agents have limited reproducibility performance, so early systems should focus on evidence extraction and reviewer assistance
  • Second, large experiments may be too expensive to rerun fully
  • Third, artifact-first requirements may disadvantage theoretical work, qualitative work, or research whose evidence is not primarily computational
  • Fourth, venue-controlled generation could introduce new biases if models are not transparent and calibrated

Open questions raised

  • Limited reproducibility performance of current AI agents (21% accuracy)
  • Scalability challenges for large/expensive experiments
  • Applicability to non-computational research (theoretical, qualitative work)
  • Transparency and calibration of venue-controlled AI generation
  • Resistance and gaming of the infrastructure by authors
  • Partial execution and remote attestation for proprietary/hardware-dependent work
Data: Motivational data: data/motivation_data.csv (secondary data for Table 5, not original experimental results); data/motivation_data.csv (for motivational data only, not original experimental results); data/motivation_data.csv (secondary data for motivation; not original experimental results)Extracted from: pdfAgreement 62%

Explore related topics

Related papers