Benchmarking LLM Agents on Meta-Analysis Articles from Nature Portfolio
Anzhe Xie, Weihang Su, Yujia Zhou, Yiqun Liu, Qingyao Ai · ArXiv.org · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
Methodology & findings
Study design
Benchmarking study with dataset construction, human annotation, system evaluation, and metric validation.
Sample
N = 442, 5 groups
Primary method
Paired t-tests (paired t with Bonferroni-style family-wise error control indicated by p<0.001 thresholds across multiple comparisons). Spearman rank correlation (ρ) for validating evaluator-dependent metrics against human judgments. Set-level metrics computed by exact ID comparison. LLM-assessed semantic similarity using an evaluator model. Pairwise agreement analysis with five-point scale collapsed to {A, tie, B} for agreement computation. Sample standard deviations available in supplement (referenced but not detailed in main text).
Main result
The study found that "no pipeline recovers more than 52.7% of ground-truth included literature, despite a retrieval ceiling of 90.9% Recall@200." The research demonstrates that "screening as the binding constraint" exists, with "the 38-point gap" between retrieval ceiling and actual pipeline performance "trace[ing] to pools averaging 16 positives among 184 PI/ECO-failing distractors, a ratio current LLMs cannot reliably resolve." Four distinct system profiles emerged: "DeepSeek-R1 produces concise analytical reports citing fewer than five studies per query"; "GLM-5 takes the opposite approach, accepting a large share of the retrieved pool to reach the highest Inc.R (up to 52.7%)"; "GPT-5 occupies a middle range"; and "ProtoMA is the precision end of the spectrum."
Reports effect sizes and confidence intervals.
Research paradigm
Empirical-computational; benchmarking and evaluation of AI systems on structured NLP tasks
Author conclusions
The authors conclude: "MetaSyn provides 442 expert-curated meta-analyses from the Nature Portfolio, each paired with a 140,585-article retrieval corpus, structured PI/ECO annotations, and ground-truth study lists." They further state: "Benchmarking twelve configurations reveals screening as the binding constraint: despite a retrieval ceiling of 90.9% at K=200, no pipeline recovers more than 52.7% of ground-truth included studies." They identify three future directions: "retrievers that exploit MetaSyn's eligibility-labeled pairs beyond a single fine-tuning pass; screening components that apply PI/ECO criteria to semantically diverse candidate pools; and synthesis metrics that accommodate certainty-aware direction claims PRISMA requires."
Risk of bias
Corpus coverage bias: PubMed-anchored corpus underrepresents EMBASE, Cochrane Central, trial registries, and grey literature sources actually used in meta-analyses; Domain-specific retrieval bias: Clinical meta-analyses in oncology draw more from specialty databases, depressing PubMed corpus match rates; non-clinical meta-analyses have higher match rates; Metric bias: Dir.A metric penalizes uncertainty-hedged conclusions, systematically disadvantaging protocol-driven systems like ProtoMA; Annotator selection: Approximately 50 annotators recruited but no diversity information provided; LLM extraction bias: PI/ECO components extracted by GLM-4.6 with human correction, but LLM biases may influence initial parsing; PubMed-only corpus may underrepresent studies indexed in EMBASE, Cochrane Central, trial registries, and grey literature; Potential annotator bias in study selection and ground truth extraction despite two-tier review; Domain-specific effects: clinical meta-analyses (67.6% of dataset) versus non-clinical (32.4%) may introduce differential evaluation bias; Database-coverage gap creates systematic underestimation of retrieval performance on clinical studies; Evaluator-model bias in semantic consistency assessment for criteria and insights metrics; Selection bias in source paper selection: restricted to Nature Portfolio journals (34,375 candidates after excluding Scientific Reports); Annotation bias: human annotators performed manual review; quality rested on annotation; ambiguous cases required adjudication; LLM-assisted extraction bias: GLM-4.6 used for initial PI/ECO parsing with human correction, but LLM errors may propagate; Database bias: PubMed-anchored corpus misses EMBASE, Cochrane Central, registries, and grey literature included in original meta-analyses; Metric bias: categorical Dir.A metric systematically penalizes protocol-driven systems with uncertainty hedging; Corpus composition bias: hard negatives (131,911 articles) selected from topically similar but ineligible papers; possible selection strategy bias in negative sampling
Limitations
- The authors state: "First, the retrieval corpus is anchored to PubMed, the largest single indexed source that systematic reviewers consult
- Because meta-analyses also draw from EMBASE, Cochrane Central, clinical-trial registries, and grey literature, on average 17.4 of a paper's 38.1 reported included studies are recovered as corpus-matched ground truth." Additionally, "Second, as documented in Section 5, the categorical Dir.A metric penalizes reports whose conclusions are phrased with explicit uncertainty, so ProtoMA's protocol-driven synthesis is systematically scored lower on this dimension than less-hedged RAG outputs."
Open questions raised
- Extending retrieval corpus to additional databases beyond PubMed (EMBASE, Cochrane Central, trial registries, grey literature)
- Retrievers that exploit MetaSyn's eligibility-labeled pairs beyond single fine-tuning pass
- Screening components that apply PI/ECO criteria to semantically diverse candidate pools
- Synthesis metrics that accommodate certainty-aware direction claims required by PRISMA
- Domain-stratified evaluation across clinical and non-clinical domains
- Certainty-aware variant of the categorical Dir.A metric
Explore related topics
Related papers
- PRISMA Extension for Scoping Reviews (PRISMA-ScR): Checklist and ExplanationAndrea C. Tricco · 2018 · 40,391 citations
- Rayyan—a web and mobile app for systematic reviewsMourad Ouzzani · 2016 · 24,664 citations
- Cochrane Handbook for Systematic Reviews of Interventions2019 · 14,420 citations
- Guidance for conducting systematic scoping reviewsMicah D.J. Peters · 2015 · 7,472 citations
- Updated methodological guidance for the conduct of scoping reviewsMicah D.J. Peters · 2020 · 6,688 citations
- Which academic search systems are suitable for systematic reviews or meta‐analyses? Evaluating retrieval qualities of Google Scholar, PubMed, and 26 other resourcesMichael Gusenbauer · 2019 · 2,116 citations