12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Jagged AI in Scientific Peer Review: Evidence from POMP Data Analysis

Jin Wook Lee, William Szegda, Zhisheng Song, Edward L. Ionides · ArXiv.org · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
2/4
Quality (LMQS)
E
Evidence
0
Citations
0.00
FWCI

Methodology & findings

Study design

Comparative empirical evaluation of four Claude AI agents versus human peer reviewers.

Sample

N = 72, 12 groups

Primary method

F-test with null hypothesis that average Human Overlap for each agent is the same, against the alternative that at least one agent's average differs. Human Overlap calculated as a proportion: (B + D + A + C) / (B + D + A + C + E), excluding category F contradictions. Thematic coding by Claude with manual verification for W21. Qualitative categorization of 411 E-category findings into five themes with frequency counts and percentages.

Main result

The study found that "AI reviewers exhibited a jagged capability profile: they proficiently caught technical implementation errors and invalid inference methodology that human reviewers often overlooked, while consistently failing to match human standards in statistical interpretation, narrative coherence, model improvement direction, and domain-informed critique." Across all four semesters and all agents, "fewer than 40% of human-confirmed weaknesses were independently identified by any AI reviewer," with the overall mean Human Overlap rates at 29.0% (Baseline), 33.4% (531-References), 31.6% (Meta-Skill), and 31.6% (Orchestrator).

Reports effect sizes.

Research paradigm

Empirical-qualitative mixed methods with quantitative assessment

Author conclusions

"The central finding was that AI reviewers exhibited a jagged capability profile: they proficiently caught technical implementation errors and invalid inference methodology that human reviewers often overlooked, while consistently failing to match human standards in statistical interpretation, narrative coherence, model improvement direction, and domain-informed critique. Skill files tuned rather than resolved this jaggedness... The complementarity between AI and human review-each detecting a largely non-overlapping set of problems-suggests that AI review can currently serve as a valuable supplement to human review, but not as a replacement for human judgement."

Risk of bias

Selection bias: Projects drawn only from a single university course (STATS 531) may not be representative of broader scientific peer review contexts; Temporal bias: Increasing human reviewer thoroughness over time (average human issues grew from 5.2 in W21 to 9.4 in W24) confounds agent performance trends; Validation bias: Human reviews were validated only by course instructor, not independently verified; Scope bias: Limited to technical POMP domain; generalizability to other scientific disciplines unclear; Operator bias: only one university course context studied; Domain-specific limitation: findings may not generalize beyond POMP analysis; Validation bias: human reviews used as gold standard without independent verification of accuracy; Temporal effect: projects from different semesters with increasing human reviewer thoroughness over time; Selection bias: only open-access projects with available source code included; Selection bias: Projects were self-selected from a single graduate course; not representative of broader scientific review contexts; Instructor validation bias: Human reviews were collated and summarized by instructors, who also added their own feedback, potentially biasing what was recorded as 'human consensus'; Temporal confounding: Human reviewer thoroughness increased across semesters (average issues grew from 5.2 to 9.4), confounding semester effects with reviewer expertise; Blinding: AI agents reviewed projects with full code access while human reviewers had the same access, but reviewer identity and potential conflicts were not mentioned; Comparison agent bias: The independent comparison agent (itself an AI system) classified findings into categories; potential algorithmic bias in categorization; Domain specificity: All projects involved POMP methodology; generalizability to other statistical domains unknown

Open questions raised

  • Performance of AI on highly technical, domain-specific statistical analyses remains largely unexplored
  • Need for multi-agent systems with independent agents assigned to specific domains (implementation correctness, inference methodology, scientific argumentation)
  • Expansion of AI focus beyond shifting between skill files without expanding total coverage
  • Journal policies regarding AI's role in peer review need development as AI-assisted analysis becomes more prevalent
  • How multi-agent systems with specialized focus areas might improve coverage across technical and scientific dimensions
  • Whether explicit instructions to cover all checks within skill files could expand focus rather than merely shifting it
Data: "All project reports are available online, anonymously, together with source code." Complete experimental results stated as "to be posted on submission."; POMP analysis projects; 72 POMP analysis projects from STATS 531 course: available online, anonymously (URL not provided in text); Complete experimental results stated as "available online (to be posted on submission)" at the time of writingCode: https://github.com/ionides/jagged-ai-peer-review (stated as location where "actual skill files and the results are available"); jagged-ai-peer-review; GitHub repository: https://github.com/ionides/jagged-ai-peer-review (containing skill files and results)Extracted from: pdfAgreement 41%

Explore related topics

Related papers