MedMeta: A Benchmark for LLMs in Synthesizing Meta-Analysis Conclusion from Medical Studies
Huy Hoang Ha, Benoit Favre, François Portet · ArXiv.org · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
Methodology & findings
Study design
Benchmark development and empirical evaluation.
Sample
N = 81, 11 groups
Primary method
Pearson correlation for linear relationships; Bland-Altman analysis for absolute agreement and systematic bias; paired t-tests (Student's t-test) to evaluate statistical significance of performance differences (α=0.05); Cohen's d for effect size estimation; Mean bias and 95% limits of agreement (LoA) calculations; Latin Square design for human annotation bias reduction; Hybrid search combining BM25 and BGE-m3 for retrieval; BERTScore for semantic similarity evaluation; LangGraph for workflow orchestration; vLLM for model inference optimization.
Main result
The study found that "RAG-based workflows consistently outperform parametric approaches across all models, establishing access to evidence as the most critical factor for high-quality synthesis." Additionally, "all models, regardless of size or specialization, fail our adversarial test by uncritically incorporating misinformation into outputs that are coherent but factually false." The research demonstrates strong correlation between LLM-as-judge and human expert evaluation (r up to 0.81), and reveals that "the benefits of domain-specific fine-tuning are likely modest and context-dependent."
Reports effect sizes and confidence intervals.
Research paradigm
Empiricist/Positivist
Author conclusions
The authors conclude: "In this work, we introduced MedMeta benchmark to evaluate the critical yet under-explored capability of multi-source conclusion synthesis in medicine. We successfully validated an LLM-J protocol, demonstrating strong alignment with experts and establishing it as a reliable, scalable proxy for evaluating medical conclusions. Our findings reveal a clear hierarchy of importance. Information grounding (RAG) provides a larger performance uplift than domain-specific fine-tuning." They further state that "current LLMs universally fail to reject factually incorrect evidence" and recommend "expanding to full-text synthesis to capture study nuances, performing multilingual evaluations to assess cross-linguistic synthesis capabilities, and building models with stronger critical reasoning to resist incorrect factual."
Risk of bias
Feasibility filtering bias: Using Gemini Flash 2.5 to assess whether conclusions could be inferred from abstracts may introduce systematic bias in which studies are retained; Sample selection bias: Reducing candidate pool from 82,233 meta-analyses to 81 through multiple filtering stages; may not be representative; Information loss bias: Using abstracts as proxy for full-text articles may systematically exclude studies requiring full-text context; Annotation bias: Limited to 9 human annotators with medical backgrounds; potential training effects and individual rater idiosyncrasies; Model selection bias: Focused hypothesis testing (H3) on single Gemma/MedGemma pair rather than all models tested; Temporal bias: Dataset restricted to 2018-2025 publications, potentially biasing against pre-2018 research methods and results; LLM-as-judge bias: Using frontier models as judges may reflect their specific biases rather than true quality assessment; LLM-based feasibility filtering may introduce bias in dataset selection; Human annotator selection bias (limited to 9 medical experts); Potential memorization bias from training data (partially mitigated by 2018-2025 publication window); Selection bias in meta-analysis curation through multi-stage filtering; Temperature setting in LLM-J evaluation (fixed at 0.0) may not capture model variance; LLM feasibility filtering bias - using Gemini Flash 2.5 for feasibility filtering may systematically bias included studies; Selection bias - only meta-analyses with structured characteristics tables and explicit conclusions were included; Limited human expert validation - only 9 annotators focusing on subset of models and conditions; Potential data contamination - while papers are from 2018-2025, authors acknowledge models may retain some pre-training knowledge; Latin Square design mitigates annotation bias but limited sample size (20 meta-analyses for validation)
Limitations
- The authors state: "First, using abstracts as a proxy for full-texts may miss study nuances
- Second, human validation was limited to 9 expert annotators and focused on a subset of models and settings
- Third, expanding beyond 81 meta-analyses and 24 specialties would further enhance its comprehensiveness
- Finally, as LLMs evolve rapidly, future work should extend this analysis to novel architectures (MoE) and paradigms (Agents)." Additionally, they acknowledge that "using an LLM to filter for feasibility could pose a risk of bias."
Open questions raised
- Multi-source conclusion synthesis capability: Identified gap that "current benchmarks do not focus on evaluating the core cognitive skill of multi-source conclusion synthesis: the ability to analyze findings from multiple, often heterogeneous, primary research articles to construct a coherent, evidence-based conclusion."
- Critical reasoning under misinformation: Gap in evaluating LLM robustness against faulty evidence
- Full-text synthesis: Need to expand beyond abstracts to capture study nuances
- Multilingual evaluation: Cross-linguistic synthesis capabilities not yet assessed
- Emerging architectures: Need for analysis on novel architectures (MoE) and paradigms (Agents)
- Current benchmarks do not evaluate multi-source conclusion synthesis - the core cognitive skill of evidence-based medicine
Explore related topics
Related papers
- PRISMA Extension for Scoping Reviews (PRISMA-ScR): Checklist and ExplanationAndrea C. Tricco · 2018 · 40,391 citations
- Cochrane Handbook for Systematic Reviews of Interventions2019 · 14,420 citations
- Updated methodological guidance for the conduct of scoping reviewsMicah D.J. Peters · 2020 · 6,688 citations
- Which academic search systems are suitable for systematic reviews or meta‐analyses? Evaluating retrieval qualities of Google Scholar, PubMed, and 26 other resourcesMichael Gusenbauer · 2019 · 2,116 citations
- Role of AI chatbots in education: systematic literature reviewLasha Labadze · 2023 · 791 citations
- Artificial intelligence adoption in the physical sciences, natural sciences, life sciences, social sciences and the arts and humanities: A bibliometric analysis of research publications from 1960-2021Stefan Hajkowicz · 2023 · 119 citations