Can AI Agents Synthesize Scientific Conclusions?
Hayoung Jung, Pedro Viana Diniz, José Reinaldo Corrêa Roveda, Abner Fernandes da Silva, Haeun Jung, Enoch Tsai et al. · ArXiv.org · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
Methodology & findings
Study design
Benchmark evaluation with controlled clean-room testing harness.
Primary method
Design science research with expert validation and iterative refinement. Includes live benchmark design with continuous updates, clean-room evaluation harness design, and expert-validated evaluation pipeline.
Main result
The study found that "the best agent achieves only a factual F1 of 0.337" under clean-room settings, and "clean-room evaluation consistently reduces performance relative to unconstrained evaluation, suggesting that leakage inflates estimates of models' true synthesis capabilities." Additionally, "44.8-84.0% of generated conclusions contained at least one fact contradicting the reference CDSR review, and nearly all contained at least one fact not supported by the reference review."
Research paradigm
positivist/empiricist
Author conclusions
The authors conclude that "reliable synthesis of scientific conclusions remains an open challenge," stating that "even under favorable settings without clean-room constraints, no system exceeds 0.63 on any metric." They emphasize that "clean-room evaluation is essential for assessing open-domain AI agents" because agents actively exploit ground-truth access, and consumer-facing systems demonstrate concerning reliability issues: "56.3% for Google AI Overview and 59% for Google AI Mode" of generated conclusions contain facts contradicting the reference reviews, raising safety concerns for high-stakes decision-making.
Risk of bias
Benchmark leakage from pre-training data; Over-filtering in clean-room protocol (false positives); Missed leakage from indirect sources (false negatives); Potential annotator bias in gold-standard dataset creation; Sample size constraints (N=268 for clean-room evaluation); Expert annotator disagreement (moderate agreement for factual precision); Data leakage from model pre-training on systematic reviews and related artifacts; Selection bias in CDSR reviews (peer-reviewed, systematic reviews may not represent all scientific evidence); Annotator expertise concentration (annotations conducted by medical students and doctors, may not represent all clinical perspectives); Over-filtering in clean-room protocol could remove legitimate content; Potential false negatives in leakage filtering allowing indirect access to reference conclusions; Selection bias: Benchmark limited to Cochrane reviews, which may not represent all scientific domains; Over-filtering risk in clean-room protocol may eliminate legitimate scientific content; Expert disagreement on factual precision judgments (moderate Cohen's κ=0.517); LLM judge may introduce bias from model-specific training data and biases; Knowledge cutoff filtering reduces sample representativeness; Missing data: Some deep research agents evaluated on reduced samples (N=100 of 268) due to cost constraints
Limitations
- The clean-room protocol "may occasionally filter legitimate content" to ensure that "the benchmark measures synthesis rather than memorization or shortcut retrieval." Additionally, the study notes that the sample size (N=268) for benchmark performance comparison "provides sufficient power to detect a statistically significant difference between two model performances in factual F1-scores of at least ∆≈0.037 at α = 0.05 with power 0.8," which may limit generalizability
- The evaluation is also limited to CDSR reviews published after model knowledge cutoffs, and only 2% of queries reached the 30 tool-call limit for two models.
Open questions raised
- Lack of evaluation frameworks for full long-horizon task of scientific conclusion synthesis from open web
- Need for live benchmarks that update to reduce benchmark leakage
- Insufficient evaluation of consumer-facing agents in high-stakes health contexts
- Limited understanding of failure modes in scientific synthesis (inverted treatment effects, mischaracterized evidence quality)
- The authors identify that prior work on AI agents "falls short in evaluating AI agents on the full long-horizon task of synthesizing long-form scientific conclusions from the open web" and that existing benchmarks "remain limited: they are often small due to the high cost of expert curation (N ≤100), become outdated as new information emerges, and fail to address benchmark leakage." They note the need for continued research on reliable synthesis and clean-room evaluation methodologies.
- Prior work focuses on intermediate artifacts (retrieval, citation grounding, summarization, short-form QA) rather than the full long-horizon task of synthesizing long-form scientific conclusions from open web
Explore related topics
Related papers
- PRISMA Extension for Scoping Reviews (PRISMA-ScR): Checklist and ExplanationAndrea C. Tricco · 2018 · 40,391 citations
- Cochrane Handbook for Systematic Reviews of Interventions2019 · 14,420 citations
- Updated methodological guidance for the conduct of scoping reviewsMicah D.J. Peters · 2020 · 6,688 citations
- Which academic search systems are suitable for systematic reviews or meta‐analyses? Evaluating retrieval qualities of Google Scholar, PubMed, and 26 other resourcesMichael Gusenbauer · 2019 · 2,116 citations
- Role of AI chatbots in education: systematic literature reviewLasha Labadze · 2023 · 791 citations
- Artificial intelligence adoption in the physical sciences, natural sciences, life sciences, social sciences and the arts and humanities: A bibliometric analysis of research publications from 1960-2021Stefan Hajkowicz · 2023 · 119 citations