12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

MedMeta: A Benchmark for LLMs in Synthesizing Meta-Analysis Conclusion from Medical Studies

Huy Hoang Ha, Benoit Favre, François Portet · ArXiv.org · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
3/4
Quality (LMQS)
E
Evidence
0
Citations
0.00
FWCI

Methodology & findings

Study design

Benchmark development and empirical evaluation.

Sample

N = 81, 11 groups

Primary method

Pearson correlation for linear relationships; Bland-Altman analysis for absolute agreement and systematic bias; paired t-tests (Student's t-test) to evaluate statistical significance of performance differences (α=0.05); Cohen's d for effect size estimation; Mean bias and 95% limits of agreement (LoA) calculations; Latin Square design for human annotation bias reduction; Hybrid search combining BM25 and BGE-m3 for retrieval; BERTScore for semantic similarity evaluation; LangGraph for workflow orchestration; vLLM for model inference optimization.

Main result

The study found that "RAG-based workflows consistently outperform parametric approaches across all models, establishing access to evidence as the most critical factor for high-quality synthesis." Additionally, "all models, regardless of size or specialization, fail our adversarial test by uncritically incorporating misinformation into outputs that are coherent but factually false." The research demonstrates strong correlation between LLM-as-judge and human expert evaluation (r up to 0.81), and reveals that "the benefits of domain-specific fine-tuning are likely modest and context-dependent."

Reports effect sizes and confidence intervals.

Research paradigm

Empiricist/Positivist

Author conclusions

The authors conclude: "In this work, we introduced MedMeta benchmark to evaluate the critical yet under-explored capability of multi-source conclusion synthesis in medicine. We successfully validated an LLM-J protocol, demonstrating strong alignment with experts and establishing it as a reliable, scalable proxy for evaluating medical conclusions. Our findings reveal a clear hierarchy of importance. Information grounding (RAG) provides a larger performance uplift than domain-specific fine-tuning." They further state that "current LLMs universally fail to reject factually incorrect evidence" and recommend "expanding to full-text synthesis to capture study nuances, performing multilingual evaluations to assess cross-linguistic synthesis capabilities, and building models with stronger critical reasoning to resist incorrect factual."

Risk of bias

Feasibility filtering bias: Using Gemini Flash 2.5 to assess whether conclusions could be inferred from abstracts may introduce systematic bias in which studies are retained; Sample selection bias: Reducing candidate pool from 82,233 meta-analyses to 81 through multiple filtering stages; may not be representative; Information loss bias: Using abstracts as proxy for full-text articles may systematically exclude studies requiring full-text context; Annotation bias: Limited to 9 human annotators with medical backgrounds; potential training effects and individual rater idiosyncrasies; Model selection bias: Focused hypothesis testing (H3) on single Gemma/MedGemma pair rather than all models tested; Temporal bias: Dataset restricted to 2018-2025 publications, potentially biasing against pre-2018 research methods and results; LLM-as-judge bias: Using frontier models as judges may reflect their specific biases rather than true quality assessment; LLM-based feasibility filtering may introduce bias in dataset selection; Human annotator selection bias (limited to 9 medical experts); Potential memorization bias from training data (partially mitigated by 2018-2025 publication window); Selection bias in meta-analysis curation through multi-stage filtering; Temperature setting in LLM-J evaluation (fixed at 0.0) may not capture model variance; LLM feasibility filtering bias - using Gemini Flash 2.5 for feasibility filtering may systematically bias included studies; Selection bias - only meta-analyses with structured characteristics tables and explicit conclusions were included; Limited human expert validation - only 9 annotators focusing on subset of models and conditions; Potential data contamination - while papers are from 2018-2025, authors acknowledge models may retain some pre-training knowledge; Latin Square design mitigates annotation bias but limited sample size (20 meta-analyses for validation)

Limitations

  • The authors state: "First, using abstracts as a proxy for full-texts may miss study nuances
  • Second, human validation was limited to 9 expert annotators and focused on a subset of models and settings
  • Third, expanding beyond 81 meta-analyses and 24 specialties would further enhance its comprehensiveness
  • Finally, as LLMs evolve rapidly, future work should extend this analysis to novel architectures (MoE) and paradigms (Agents)." Additionally, they acknowledge that "using an LLM to filter for feasibility could pose a risk of bias."

Open questions raised

  • Multi-source conclusion synthesis capability: Identified gap that "current benchmarks do not focus on evaluating the core cognitive skill of multi-source conclusion synthesis: the ability to analyze findings from multiple, often heterogeneous, primary research articles to construct a coherent, evidence-based conclusion."
  • Critical reasoning under misinformation: Gap in evaluating LLM robustness against faulty evidence
  • Full-text synthesis: Need to expand beyond abstracts to capture study nuances
  • Multilingual evaluation: Cross-linguistic synthesis capabilities not yet assessed
  • Emerging architectures: Need for analysis on novel architectures (MoE) and paradigms (Agents)
  • Current benchmarks do not evaluate multi-source conclusion synthesis - the core cognitive skill of evidence-based medicine
Data: MedMeta benchmark: 81 meta-analyses from PubMed (2018-2025) - stated as "available in our public repository" in the paper; Evaluation corpus: 25,000 PubMed abstracts (2,250 ground-truth abstracts from meta-analyses + 22,750 random noise abstracts) constructed for Standard-RAG evaluation; MedMeta benchmark (81 meta-analyses, 2,250 primary studies, 24 medical specialties) - stated as 'available in our public repository'; 25,000 PubMed abstracts (2,250 ground-truth + 22,750 noise abstracts) used for Standard-RAG evaluation; MedMeta benchmark: 81 curated meta-analyses from PubMed (2018-2025) with 2,250 primary studies; Proxy evaluation corpus: 25,000 PubMed abstracts (2,250 ground-truth + 22,750 noise abstracts); Authors state: "The complete benchmark is available in our public repository" and "will release the source code and scripts in an anonymous repository during the review process"Code: Anonymous repository to be released during review process containing: Data folder with preprocessed MedMeta dataset, Scripts folder with step-by-step reproduction scripts, Src directory with LangGraph implementations, and Web folder with annotation platform source code; Anonymous repository to be released during review process containing: (i) Data folder with preprocessed MedMeta dataset, (ii) Scripts folder with step-by-step reproduction scripts, (iii) Src directory with LangGraph implementations, (iv) Web folder with annotation platform source code, and comprehensive README with setup instructions.; Anonymous repository to be released during review containing: (i) Data folder with MedMeta dataset, (ii) Scripts folder with reproduction scripts, (iii) Src directory with LangGraph implementations, (iv) Web folder with annotation platform source codeExtracted from: pdfAgreement 54%

Explore related topics

Related papers