Jagged AI in Scientific Peer Review: Evidence from POMP Data Analysis
Jin Wook Lee, William Szegda, Zhisheng Song, Edward L. Ionides · ArXiv.org · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
Methodology & findings
Study design
Comparative empirical evaluation of four Claude AI agents versus human peer reviewers.
Sample
N = 72, 12 groups
Primary method
F-test with null hypothesis that average Human Overlap for each agent is the same, against the alternative that at least one agent's average differs. Human Overlap calculated as a proportion: (B + D + A + C) / (B + D + A + C + E), excluding category F contradictions. Thematic coding by Claude with manual verification for W21. Qualitative categorization of 411 E-category findings into five themes with frequency counts and percentages.
Main result
The study found that "AI reviewers exhibited a jagged capability profile: they proficiently caught technical implementation errors and invalid inference methodology that human reviewers often overlooked, while consistently failing to match human standards in statistical interpretation, narrative coherence, model improvement direction, and domain-informed critique." Across all four semesters and all agents, "fewer than 40% of human-confirmed weaknesses were independently identified by any AI reviewer," with the overall mean Human Overlap rates at 29.0% (Baseline), 33.4% (531-References), 31.6% (Meta-Skill), and 31.6% (Orchestrator).
Reports effect sizes.
Research paradigm
Empirical-qualitative mixed methods with quantitative assessment
Author conclusions
"The central finding was that AI reviewers exhibited a jagged capability profile: they proficiently caught technical implementation errors and invalid inference methodology that human reviewers often overlooked, while consistently failing to match human standards in statistical interpretation, narrative coherence, model improvement direction, and domain-informed critique. Skill files tuned rather than resolved this jaggedness... The complementarity between AI and human review-each detecting a largely non-overlapping set of problems-suggests that AI review can currently serve as a valuable supplement to human review, but not as a replacement for human judgement."
Risk of bias
Selection bias: Projects drawn only from a single university course (STATS 531) may not be representative of broader scientific peer review contexts; Temporal bias: Increasing human reviewer thoroughness over time (average human issues grew from 5.2 in W21 to 9.4 in W24) confounds agent performance trends; Validation bias: Human reviews were validated only by course instructor, not independently verified; Scope bias: Limited to technical POMP domain; generalizability to other scientific disciplines unclear; Operator bias: only one university course context studied; Domain-specific limitation: findings may not generalize beyond POMP analysis; Validation bias: human reviews used as gold standard without independent verification of accuracy; Temporal effect: projects from different semesters with increasing human reviewer thoroughness over time; Selection bias: only open-access projects with available source code included; Selection bias: Projects were self-selected from a single graduate course; not representative of broader scientific review contexts; Instructor validation bias: Human reviews were collated and summarized by instructors, who also added their own feedback, potentially biasing what was recorded as 'human consensus'; Temporal confounding: Human reviewer thoroughness increased across semesters (average issues grew from 5.2 to 9.4), confounding semester effects with reviewer expertise; Blinding: AI agents reviewed projects with full code access while human reviewers had the same access, but reviewer identity and potential conflicts were not mentioned; Comparison agent bias: The independent comparison agent (itself an AI system) classified findings into categories; potential algorithmic bias in categorization; Domain specificity: All projects involved POMP methodology; generalizability to other statistical domains unknown
Open questions raised
- Performance of AI on highly technical, domain-specific statistical analyses remains largely unexplored
- Need for multi-agent systems with independent agents assigned to specific domains (implementation correctness, inference methodology, scientific argumentation)
- Expansion of AI focus beyond shifting between skill files without expanding total coverage
- Journal policies regarding AI's role in peer review need development as AI-assisted analysis becomes more prevalent
- How multi-agent systems with specialized focus areas might improve coverage across technical and scientific dimensions
- Whether explicit instructions to cover all checks within skill files could expand focus rather than merely shifting it
Explore related topics
Related papers
- Rayyan—a web and mobile app for systematic reviewsMourad Ouzzani · 2016 · 24,664 citations
- What if the devil is my guardian angel: ChatGPT as a case study of using chatbots in educationAhmed Tlili · 2023 · 1,587 citations
- Embracing the future of Artificial Intelligence in the classroom: the relevance of AI literacy, prompt engineering, and critical thinking in modern educationYoshija Walter · 2024 · 805 citations
- Generative AI tools and assessment: Guidelines of the world's top-ranking universitiesBenjamin Luke Moorhouse · 2023 · 343 citations
- AI-assisted peer reviewAlessandro Checco · 2021 · 261 citations
- Leveraging ChatGPT for Enhancing Critical Thinking SkillsYing Guo · 2023 · 223 citations