12,637 papers · updated 18 Sept 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Jagged AI in Scientific Peer Review: Evidence from POMP Data Analysis

Jin Wook Lee, William Szegda, Zhisheng Song, Edward L. Ionides · ArXiv.org · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
E
Evidence
0
Citations
0.00
FWCI

Methodology & findings

Study design

Comparative empirical evaluation of four Claude AI agents versus human peer reviewers.

Sample

N = 72, 4 groups

Primary method

F-test with null hypothesis that average Human Overlap for each agent is the same, against the alternative that at least one agent's average differs. Human Overlap calculated as a proportion: (B + D + A + C) / (B + D + A + C + E), excluding category F contradictions. Thematic coding by Claude with manual verification for W21. Qualitative categorization of 411 E-category findings into five themes with frequency counts and percentages.

Main result

The study found that "AI reviewers exhibited a jagged capability profile: they proficiently caught technical implementation errors and invalid inference methodology that human reviewers often overlooked, while consistently failing to match human standards in statistical interpretation, narrative coherence, model improvement direction, and domain-informed critique." Across all four semesters and all agents, "fewer than 40% of human-confirmed weaknesses were independently identified by any AI reviewer," with the overall mean Human Overlap rates at 29.0% (Baseline), 33.4% (531-References), 31.6% (Meta-Skill), and 31.6% (Orchestrator).

Reports effect sizes.

Research paradigm

Empirical-qualitative mixed methods with quantitative assessment

Author conclusions

"The central finding was that AI reviewers exhibited a jagged capability profile: they proficiently caught technical implementation errors and invalid inference methodology that human reviewers often overlooked, while consistently failing to match human standards in statistical interpretation, narrative coherence, model improvement direction, and domain-informed critique. Skill files tuned rather than resolved this jaggedness... The complementarity between AI and human review-each detecting a largely non-overlapping set of problems-suggests that AI review can currently serve as a valuable supplement to human review, but not as a replacement for human judgement."

Risk of bias

Selection bias: Projects drawn only from a single university course (STATS 531) may not be representative of broader scientific peer review contexts; Temporal bias: Increasing human reviewer thoroughness over time (average human issues grew from 5.2 in W21 to 9.4 in W24) confounds agent performance trends; Validation bias: Human reviews were validated only by course instructor, not independently verified; Scope bias: Limited to technical POMP domain; generalizability to other scientific disciplines unclear; Domain-specific limitation: findings may not generalize beyond POMP analysis; Selection bias: only open-access projects with available source code included; Instructor validation bias: Human reviews were collated and summarized by instructors, who also added their own feedback, potentially biasing what was recorded as 'human consensus'; Blinding: AI agents reviewed projects with full code access while human reviewers had the same access, but reviewer identity and potential conflicts were not mentioned; Comparison agent bias: The independent comparison agent (itself an AI system) classified findings into categories; potential algorithmic bias in categorization; Domain specificity: All projects involved POMP methodology; generalizability to other statistical domains unknown

Open questions raised

  • Performance of AI on highly technical, domain-specific statistical analyses remains largely unexplored
  • Need for multi-agent systems with independent agents assigned to specific domains (implementation correctness, inference methodology, scientific argumentation)
  • Expansion of AI focus beyond shifting between skill files without expanding total coverage
  • Journal policies regarding AI's role in peer review need development as AI-assisted analysis becomes more prevalent
  • How multi-agent systems with specialized focus areas might improve coverage across technical and scientific dimensions
  • Whether explicit instructions to cover all checks within skill files could expand focus rather than merely shifting it
Data: "All project reports are available online, anonymously, together with source code." Complete experimental results stated as "to be posted on submission."; POMP analysis projectsCode: https://github.com/ionides/jagged-ai-peer-review (stated as location where "actual skill files and the results are available")Extracted from: pdf

Explore related topics

Related papers