12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Benchmarking LLM Agents on Meta-Analysis Articles from Nature Portfolio

Anzhe Xie, Weihang Su, Yujia Zhou, Yiqun Liu, Qingyao Ai · ArXiv.org · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
2/4
Quality (LMQS)
E
Evidence
0
Citations
0.00
FWCI

Methodology & findings

Study design

Benchmarking study with dataset construction, human annotation, system evaluation, and metric validation.

Sample

N = 442, 5 groups

Primary method

Paired t-tests (paired t with Bonferroni-style family-wise error control indicated by p<0.001 thresholds across multiple comparisons). Spearman rank correlation (ρ) for validating evaluator-dependent metrics against human judgments. Set-level metrics computed by exact ID comparison. LLM-assessed semantic similarity using an evaluator model. Pairwise agreement analysis with five-point scale collapsed to {A, tie, B} for agreement computation. Sample standard deviations available in supplement (referenced but not detailed in main text).

Main result

The study found that "no pipeline recovers more than 52.7% of ground-truth included literature, despite a retrieval ceiling of 90.9% Recall@200." The research demonstrates that "screening as the binding constraint" exists, with "the 38-point gap" between retrieval ceiling and actual pipeline performance "trace[ing] to pools averaging 16 positives among 184 PI/ECO-failing distractors, a ratio current LLMs cannot reliably resolve." Four distinct system profiles emerged: "DeepSeek-R1 produces concise analytical reports citing fewer than five studies per query"; "GLM-5 takes the opposite approach, accepting a large share of the retrieved pool to reach the highest Inc.R (up to 52.7%)"; "GPT-5 occupies a middle range"; and "ProtoMA is the precision end of the spectrum."

Reports effect sizes and confidence intervals.

Research paradigm

Empirical-computational; benchmarking and evaluation of AI systems on structured NLP tasks

Author conclusions

The authors conclude: "MetaSyn provides 442 expert-curated meta-analyses from the Nature Portfolio, each paired with a 140,585-article retrieval corpus, structured PI/ECO annotations, and ground-truth study lists." They further state: "Benchmarking twelve configurations reveals screening as the binding constraint: despite a retrieval ceiling of 90.9% at K=200, no pipeline recovers more than 52.7% of ground-truth included studies." They identify three future directions: "retrievers that exploit MetaSyn's eligibility-labeled pairs beyond a single fine-tuning pass; screening components that apply PI/ECO criteria to semantically diverse candidate pools; and synthesis metrics that accommodate certainty-aware direction claims PRISMA requires."

Risk of bias

Corpus coverage bias: PubMed-anchored corpus underrepresents EMBASE, Cochrane Central, trial registries, and grey literature sources actually used in meta-analyses; Domain-specific retrieval bias: Clinical meta-analyses in oncology draw more from specialty databases, depressing PubMed corpus match rates; non-clinical meta-analyses have higher match rates; Metric bias: Dir.A metric penalizes uncertainty-hedged conclusions, systematically disadvantaging protocol-driven systems like ProtoMA; Annotator selection: Approximately 50 annotators recruited but no diversity information provided; LLM extraction bias: PI/ECO components extracted by GLM-4.6 with human correction, but LLM biases may influence initial parsing; PubMed-only corpus may underrepresent studies indexed in EMBASE, Cochrane Central, trial registries, and grey literature; Potential annotator bias in study selection and ground truth extraction despite two-tier review; Domain-specific effects: clinical meta-analyses (67.6% of dataset) versus non-clinical (32.4%) may introduce differential evaluation bias; Database-coverage gap creates systematic underestimation of retrieval performance on clinical studies; Evaluator-model bias in semantic consistency assessment for criteria and insights metrics; Selection bias in source paper selection: restricted to Nature Portfolio journals (34,375 candidates after excluding Scientific Reports); Annotation bias: human annotators performed manual review; quality rested on annotation; ambiguous cases required adjudication; LLM-assisted extraction bias: GLM-4.6 used for initial PI/ECO parsing with human correction, but LLM errors may propagate; Database bias: PubMed-anchored corpus misses EMBASE, Cochrane Central, registries, and grey literature included in original meta-analyses; Metric bias: categorical Dir.A metric systematically penalizes protocol-driven systems with uncertainty hedging; Corpus composition bias: hard negatives (131,911 articles) selected from topically similar but ineligible papers; possible selection strategy bias in negative sampling

Limitations

  • The authors state: "First, the retrieval corpus is anchored to PubMed, the largest single indexed source that systematic reviewers consult
  • Because meta-analyses also draw from EMBASE, Cochrane Central, clinical-trial registries, and grey literature, on average 17.4 of a paper's 38.1 reported included studies are recovered as corpus-matched ground truth." Additionally, "Second, as documented in Section 5, the categorical Dir.A metric penalizes reports whose conclusions are phrased with explicit uncertainty, so ProtoMA's protocol-driven synthesis is systematically scored lower on this dimension than less-hedged RAG outputs."

Open questions raised

  • Extending retrieval corpus to additional databases beyond PubMed (EMBASE, Cochrane Central, trial registries, grey literature)
  • Retrievers that exploit MetaSyn's eligibility-labeled pairs beyond single fine-tuning pass
  • Screening components that apply PI/ECO criteria to semantically diverse candidate pools
  • Synthesis metrics that accommodate certainty-aware direction claims required by PRISMA
  • Domain-stratified evaluation across clinical and non-clinical domains
  • Certainty-aware variant of the categorical Dir.A metric
Data: MetaSyn dataset: 442 expert-curated meta-analyses from Nature Portfolio journals with 140,585-article retrieval corpus (access method not specified in paper); MetaSyn: 442 meta-analyses from Nature Portfolio with 140,585-article retrieval corpus (availability status not explicitly stated in paper); MetaSyn dataset: 442 expert-curated meta-analyses from Nature Portfolio journals with 140,585-article retrieval corpus (availability status not explicitly stated in paper)Extracted from: pdfAgreement 58%

Explore related topics

Related papers