12,637 papers · updated 18 Sept 2026livingmeta.ai
← Browse all papers
AI evidence extraction

MedResearchBench: A Multi-Domain Benchmark for Evaluating AI Research Agents on Clinical Medical Research

Shuping Tan, Zhanxiao Tian · medRxiv · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
E
Evidence
0
Citations
0.00
FWCI

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.64898/2026.03.30.26349749

Methodology & findings

Study design

Benchmark design and construction with pilot validation.

Sample

N = 3, 3 groups

Primary method

Weighted scoring methodology using LLM Judge evaluation. Objective mode: each checklist item scored 0-100 where 50 represents matching ground truth paper quality, with weighted sum across items producing final task score. Dimension-specific weights applied: statistical methodology (w=0.20), results accuracy (w=0.25), visualization quality (w=0.15), clinical interpretation (w=0.20), confounding sensitivity (w=0.10), reporting compliance (w=0.10). Supplementary subjective mode assessment of overall research quality, novelty, and clinical utility. Medical-specific extensions include survey design penalty and STROBE compliance checklist. Evaluation performed using GPT-4o or Claude Opus LLM Judge with prompt templates provided.

Main result

Initial baseline evaluation demonstrated that "End-to-end evaluation of an agentic pipeline across 3 pilot tasks (Tier 1-3), yielding a mean score of 72/100 (B-level), establishing the first quantitative baseline for AI-driven medical research quality on this benchmark." The pilot results showed survey-weighted methodology was correctly implemented with 100% compliance across all three tasks, but results accuracy was the primary limitation with a mean score of 58.7/100, driven by covariate incompleteness and reference group misspecification.

Reports effect sizes and confidence intervals.

Research paradigm

Positivist/empiricist - quantitative evaluation through systematic benchmarking

Author conclusions

The authors conclude: "As AI systems increasingly automate scientific research, the medical domain-with its high stakes, complex methodology, and established quality standards-demands specialized evaluation. MedResearchBench provides that evaluation platform and, through its anti-paper-mill design, serves as a quality gate for responsible AI-assisted medical research." They state the benchmark "is designed as a complement to ResearchClawBench, not a replacement. Together, the two benchmarks span the full landscape of AI research evaluation, from fundamental science to clinical medicine."

Risk of bias

Limited pilot evaluation (n=3 tasks only); Potential bias toward NHANES-based tasks (13/16 tasks); Ground truth paper selection bias: papers selected must meet specific quality criteria and use public datasets; LLM Judge evaluation dependency: scoring relies on GPT-4o or Claude Opus capabilities; Evaluation limited to analytical pipeline; does not capture full clinical research complexity; No inter-rater reliability reported for LLM Judge evaluation; Ground truth papers span wide IF range (2.3-51.0) but may not capture full distribution of real-world medical publications

Limitations

  • The authors explicitly state: "No wet-lab or clinical data collection: Tasks evaluate the analytical pipeline only." Additional limitations are implied in the task design: "visualization quality appears in only 2 tasks in Phase 1
  • Phase 2 will systematically increase visualization coverage." The benchmark is limited to Phase 1 tasks (16 tasks), and evaluation focuses on end-to-end analytical workflows from publicly available datasets rather than prospective data collection.

Open questions raised

  • No existing benchmark tests whether an AI system can correctly go from patient-level data to a clinically sound, publication-ready manuscript
  • Medical research tools remain comparatively underdeveloped compared to fundamental science automation
  • No published system specifically targets end-to-end automation of observational clinical research
  • Future work: Phase 2 will systematically increase visualization coverage (currently only 2/16 tasks)
  • Potential for extending benchmark to prospective validation with new published papers
  • No existing benchmark evaluates AI systems on medical clinical research tasks prior to MedResearchBench
Data: NHANES (National Health and Nutrition Examination Survey) - 11 tasks, cycles 1999-2000 through 2017-2018; NHANES III with NDI linkage - 2 tasks, 1988-1994 with mortality follow-up through 2019; SEER (Surveillance, Epidemiology, and End Results) - 2 oncology tasks, 1991-2018; SEER-Medicare - oncology task Onco_002; Benchmark repository: https://github.com/TerryFYL/MedResearchBench/tree/main/bench-runs (pilot evaluation reports available here)Code: https://github.com/TerryFYL/MedResearchBench/tree/main/bench-runs (detailed evaluation reports)Extracted from: pdf

Explore related topics

Related papers