ARA: Agentic Reproducibility Assessment For Scalable Support Of Scientific Peer-Review
Kevin Riehl, Andrés L. Marín, Nikofors Zacharof, Fan Wu, Patrick Langer, Robert Jakob et al. · ArXiv.org · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
Methodology & findings
Study design
Empirical evaluation study using LLM-based agentic pipeline to extract and score scientific workflows from published papers.
Sample
N = 375, 8 groups
Primary method
Graph edit distance (GED) - normalized mean pairwise labeled edit distance for workflow graph comparison; Standard deviation computation across repeated runs for stability assessment; Accuracy metrics (ACC) comparing ARA scores to human annotations; F1 scores for multi-class classification performance; Score distance and absolute score distance metrics for ordinal agreement; Signed and unsigned disagreement analysis by workflow stage and document characteristics; Aggregate statistics across paper-model-temperature configurations
Main result
The study found that "structured workflow reconstruction enables agentic systems to approximate human reproducibility judgments with moderate agreement despite operating without access to external artifacts, replication packages, or web-based information sources." Across the ReScience C corpus of 213 papers, ARA achieved "accuracies around 60% despite operating strictly under a document-level evaluation setting," with "agreement highest for sources and sinks, where reporting structures are typically explicit, and lower for methods and experiments, which depend more strongly on implementation details revealed only during reproduction."
Reports effect sizes and confidence intervals.
Research paradigm
Empiricist/Positivist - computational reproducibility assessment through structured reasoning and workflow reconstruction
Author conclusions
The authors conclude that "agentic, document-level assessments could serve as a scalable diagnostic complement, not a replacement of human expert peer review" and that this "creates opportunities for the acceleration of editorial pipelines, large-scale literature screening systems, and meta-research infrastructure supporting transparency and reliability in scientific communication." They emphasize that "acceptance of such assistance systems will, as we argue, largely depend on transparency."
Risk of bias
Selection bias: ReScience C papers are from a journal dedicated to reproducibility, potentially representing papers more amenable to reproduction than the general scientific literature; Information asymmetry: Human reference labels derived from full replication efforts may incorporate execution-level details not observable in text alone; Model-dependent variation: Different LLM architectures and temperatures produce variable outputs; some configurations exhibit higher failure rates; Document length bias: Shorter papers show lower reproducibility scores overall; medium-length papers exhibit highest disagreement between agent and human assessments; Domain-specific reporting norms: Agreement varies by workflow stage (highest for sources/sinks, lower for methods/experiments); Document-level assessment bias: ARA cannot detect implementation barriers visible only during execution; Selection bias in benchmark datasets: ReScience C, ReproBench, and GoldStandardDB represent pre-selected reproducible or attempted replications, not random scientific literature; Reference standard bias: Human annotations derived from full replication efforts represent gold standard but incorporate external information unavailable to ARA; Model-temperature variability: Stochastic inference may introduce inconsistency across LLM runs; Domain heterogeneity: Multi-domain evaluation may obscure domain-specific performance differences; Selection bias: ReScience C papers are from a specialized reproducibility-focused journal, not representative of all scientific publications; Document-level assessment limitations: ARA cannot capture implementation barriers discoverable only during full replication; Context window constraints: Local models unable to process papers exceeding ~15,000 words; Potential domain bias: Journal-specific reporting practices may not generalize across all scientific fields
Limitations
- "Human reproducibility assessments often incorporate external information sources such as repositories, search engines, and author communication, whereas ARA operates strictly at the document level
- Moreover, the human reference labels (used here) are derived from full replication efforts that involve implementing methods, executing experiments, and validating reported results over extended periods of time
- These processes can reveal practical reproducibility barriers that are not observable from the publication text alone, making direct comparisons between document-level assessments and execution-based replication outcomes inherently conservative for ARA." Additionally, "local models have more limited context capacity than commercial systems, especially for documents exceeding approximately 15,000 words," and some papers exceeded context windows of reduced models.
Open questions raised
- Future work needed on improving reproducibility assessment for methods and experiments stages, where human-agent disagreement is highest due to implementation details revealed only during replication
- Need for domain-specific refinement of framework despite demonstrated domain-agnostic capabilities
- Scaling to broader publication corpora beyond ReScience C
- Integration with extended scientific workflows and execution environments
- Governance frameworks for transparency in AI-assisted peer review
- Need for automation of reproducibility assessment at scale to address reviewer workload and publication volume growth
Explore related topics
Related papers
- Estimating the reproducibility of psychological scienceAlexander A. Aarts · 2015 · 8,669 citations
- Which academic search systems are suitable for systematic reviews or meta‐analyses? Evaluating retrieval qualities of Google Scholar, PubMed, and 26 other resourcesMichael Gusenbauer · 2019 · 2,116 citations
- Generative AI tools and assessment: Guidelines of the world's top-ranking universitiesBenjamin Luke Moorhouse · 2023 · 343 citations
- AI-assisted peer reviewAlessandro Checco · 2021 · 261 citations
- Fighting reviewer fatigue or amplifying bias? Considerations and recommendations for use of ChatGPT and other large language models in scholarly peer reviewMohammad Hosseini · 2023 · 209 citations
- Artificial intelligence to support publishing and peer review: A summary and reviewKayvan Kousha · 2023 · 137 citations