12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Can AI Agents Synthesize Scientific Conclusions?

Hayoung Jung, Pedro Viana Diniz, José Reinaldo Corrêa Roveda, Abner Fernandes da Silva, Haeun Jung, Enoch Tsai et al. · ArXiv.org · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
1/4
Quality (LMQS)
D
Evidence
0
Citations
0.00
FWCI

Methodology & findings

Study design

Benchmark evaluation with controlled clean-room testing harness.

Primary method

Design science research with expert validation and iterative refinement. Includes live benchmark design with continuous updates, clean-room evaluation harness design, and expert-validated evaluation pipeline.

Main result

The study found that "the best agent achieves only a factual F1 of 0.337" under clean-room settings, and "clean-room evaluation consistently reduces performance relative to unconstrained evaluation, suggesting that leakage inflates estimates of models' true synthesis capabilities." Additionally, "44.8-84.0% of generated conclusions contained at least one fact contradicting the reference CDSR review, and nearly all contained at least one fact not supported by the reference review."

Research paradigm

positivist/empiricist

Author conclusions

The authors conclude that "reliable synthesis of scientific conclusions remains an open challenge," stating that "even under favorable settings without clean-room constraints, no system exceeds 0.63 on any metric." They emphasize that "clean-room evaluation is essential for assessing open-domain AI agents" because agents actively exploit ground-truth access, and consumer-facing systems demonstrate concerning reliability issues: "56.3% for Google AI Overview and 59% for Google AI Mode" of generated conclusions contain facts contradicting the reference reviews, raising safety concerns for high-stakes decision-making.

Risk of bias

Benchmark leakage from pre-training data; Over-filtering in clean-room protocol (false positives); Missed leakage from indirect sources (false negatives); Potential annotator bias in gold-standard dataset creation; Sample size constraints (N=268 for clean-room evaluation); Expert annotator disagreement (moderate agreement for factual precision); Data leakage from model pre-training on systematic reviews and related artifacts; Selection bias in CDSR reviews (peer-reviewed, systematic reviews may not represent all scientific evidence); Annotator expertise concentration (annotations conducted by medical students and doctors, may not represent all clinical perspectives); Over-filtering in clean-room protocol could remove legitimate content; Potential false negatives in leakage filtering allowing indirect access to reference conclusions; Selection bias: Benchmark limited to Cochrane reviews, which may not represent all scientific domains; Over-filtering risk in clean-room protocol may eliminate legitimate scientific content; Expert disagreement on factual precision judgments (moderate Cohen's κ=0.517); LLM judge may introduce bias from model-specific training data and biases; Knowledge cutoff filtering reduces sample representativeness; Missing data: Some deep research agents evaluated on reduced samples (N=100 of 268) due to cost constraints

Limitations

  • The clean-room protocol "may occasionally filter legitimate content" to ensure that "the benchmark measures synthesis rather than memorization or shortcut retrieval." Additionally, the study notes that the sample size (N=268) for benchmark performance comparison "provides sufficient power to detect a statistically significant difference between two model performances in factual F1-scores of at least ∆≈0.037 at α = 0.05 with power 0.8," which may limit generalizability
  • The evaluation is also limited to CDSR reviews published after model knowledge cutoffs, and only 2% of queries reached the 30 tool-call limit for two models.

Open questions raised

  • Lack of evaluation frameworks for full long-horizon task of scientific conclusion synthesis from open web
  • Need for live benchmarks that update to reduce benchmark leakage
  • Insufficient evaluation of consumer-facing agents in high-stakes health contexts
  • Limited understanding of failure modes in scientific synthesis (inverted treatment effects, mischaracterized evidence quality)
  • The authors identify that prior work on AI agents "falls short in evaluating AI agents on the full long-horizon task of synthesizing long-form scientific conclusions from the open web" and that existing benchmarks "remain limited: they are often small due to the high cost of expert curation (N ≤100), become outdated as new information emerges, and fail to address benchmark leakage." They note the need for continued research on reliable synthesis and clean-room evaluation methodologies.
  • Prior work focuses on intermediate artifacts (retrieval, citation grounding, summarization, short-form QA) rather than the full long-horizon task of synthesizing long-form scientific conclusions from open web
Data: SCICONBENCH dataset: hayoungjung/SciConBench (available on Hugging Face Hub); Cochrane Database of Systematic Reviews (CDSR) - 9,107 valid reviews used; SCICONBENCH: https://huggingface.co/datasets/hayoungjung/SciConBench; Code: https://github.com/hayoungjungg/SciConBench; SCICONBENCH: Available at https://huggingface.co/datasets/hayoungjung/SciConBench (9,107 valid samples derived from Cochrane Database of Systematic Reviews)Code: https://github.com/hayoungjungg/SciConBenchExtracted from: pdfAgreement 51%

Explore related topics

Related papers