12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

MedResearchBench: A Multi-Domain Benchmark for Evaluating AI Research Agents on Clinical Medical Research

Shuping Tan, Zhanxiao Tian · medRxiv · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
1/4
Quality (LMQS)
E
Evidence
0
Citations
0.00
FWCI

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.64898/2026.03.30.26349749

Methodology & findings

Study design

Benchmark design and construction with pilot validation.

Sample

N = 3, 7 groups

Primary method

Weighted scoring methodology using LLM Judge evaluation. Objective mode: each checklist item scored 0-100 where 50 represents matching ground truth paper quality, with weighted sum across items producing final task score. Dimension-specific weights applied: statistical methodology (w=0.20), results accuracy (w=0.25), visualization quality (w=0.15), clinical interpretation (w=0.20), confounding sensitivity (w=0.10), reporting compliance (w=0.10). Supplementary subjective mode assessment of overall research quality, novelty, and clinical utility. Medical-specific extensions include survey design penalty and STROBE compliance checklist. Evaluation performed using GPT-4o or Claude Opus LLM Judge with prompt templates provided.

Main result

Initial baseline evaluation demonstrated that "End-to-end evaluation of an agentic pipeline across 3 pilot tasks (Tier 1-3), yielding a mean score of 72/100 (B-level), establishing the first quantitative baseline for AI-driven medical research quality on this benchmark." The pilot results showed survey-weighted methodology was correctly implemented with 100% compliance across all three tasks, but results accuracy was the primary limitation with a mean score of 58.7/100, driven by covariate incompleteness and reference group misspecification.

Reports effect sizes and confidence intervals.

Research paradigm

Positivist/empiricist - quantitative evaluation through systematic benchmarking

Author conclusions

The authors conclude: "As AI systems increasingly automate scientific research, the medical domain-with its high stakes, complex methodology, and established quality standards-demands specialized evaluation. MedResearchBench provides that evaluation platform and, through its anti-paper-mill design, serves as a quality gate for responsible AI-assisted medical research." They state the benchmark "is designed as a complement to ResearchClawBench, not a replacement. Together, the two benchmarks span the full landscape of AI research evaluation, from fundamental science to clinical medicine."

Risk of bias

Limited pilot evaluation (n=3 tasks only); Potential bias toward NHANES-based tasks (13/16 tasks); Ground truth paper selection bias: papers selected must meet specific quality criteria and use public datasets; LLM Judge evaluation dependency: scoring relies on GPT-4o or Claude Opus capabilities; Evaluation limited to analytical pipeline; does not capture full clinical research complexity; Selection bias in ground truth paper selection (papers must meet authenticity, quality, relevance, methodological rigor criteria and anti-paper-mill filtering). Task difficulty stratification may introduce complexity bias. LLM Judge evaluation introduces potential bias depending on model choice (GPT-4o or Claude Opus). Evaluation limited to publicly available datasets may not represent all clinical research paradigms.; Selection bias in ground truth paper selection (papers meeting specific criteria may not represent full diversity of medical research quality); LLM Judge bias (evaluation depends on GPT-4o or Claude Opus behavior, which may have inherent biases); Limited baseline validation (only 3 tasks evaluated, not the full 16-task benchmark); NHANES dominance (13/16 tasks) may introduce dataset-specific bias; No inter-rater reliability reported for LLM Judge evaluation; Ground truth papers span wide IF range (2.3-51.0) but may not capture full distribution of real-world medical publications

Limitations

  • The authors explicitly state: "No wet-lab or clinical data collection: Tasks evaluate the analytical pipeline only." Additional limitations are implied in the task design: "visualization quality appears in only 2 tasks in Phase 1
  • Phase 2 will systematically increase visualization coverage." The benchmark is limited to Phase 1 tasks (16 tasks), and evaluation focuses on end-to-end analytical workflows from publicly available datasets rather than prospective data collection.

Open questions raised

  • No existing benchmark tests whether an AI system can correctly go from patient-level data to a clinically sound, publication-ready manuscript
  • Medical research tools remain comparatively underdeveloped compared to fundamental science automation
  • No published system specifically targets end-to-end automation of observational clinical research
  • Future work: Phase 2 will systematically increase visualization coverage (currently only 2/16 tasks)
  • Potential for extending benchmark to prospective validation with new published papers
  • The authors identify that no existing benchmark tests "whether an AI system can correctly go from patient-level data to a clinically sound, publication-ready manuscript." They note that ResearchClawBench, the most comprehensive existing benchmark, "excludes clinical medicine entirely" from its 10 fundamental science domains. Medical research tools remain "comparatively underdeveloped" with "no published system specifically target[ing] the end-to-end automation of observational clinical research." Future directions mentioned include systematic increase of visualization quality coverage (Phase 2 will increase from 2 to more tasks), and extension beyond Phase 1's 16 tasks.
Data: NHANES (National Health and Nutrition Examination Survey) - 11 tasks, cycles 1999-2000 through 2017-2018; NHANES III with NDI linkage - 2 tasks, 1988-1994 with mortality follow-up through 2019; SEER (Surveillance, Epidemiology, and End Results) - 2 oncology tasks, 1991-2018; SEER-Medicare - oncology task Onco_002; NHANES (National Health and Nutrition Examination Survey) - cycles from 1999-2000 through 2017-2018, and NHANES III (1988-1994) with NDI mortality linkage; SEER (Surveillance, Epidemiology, and End Results) and SEER-Medicare registries (1991-2018); Benchmark repository: https://github.com/TerryFYL/MedResearchBench/tree/main/bench-runs (pilot evaluation reports available here); NHANES (National Health and Nutrition Examination Survey) - 1999-2000 through 2017-2018 cycles; NHANES III - 1988-1994 with NDI mortality linkage through 2019; SEER (Surveillance, Epidemiology, and End Results) - 1991-2018; SEER-Medicare - 1991-2018; Repository: https://github.com/TerryFYL/MedResearchBench/Code: https://github.com/TerryFYL/MedResearchBench/tree/main/bench-runs (detailed evaluation reports); https://github.com/TerryFYL/MedResearchBench (main benchmark repository); https://github.com/TerryFYL/MedResearchBench/ (includes task definitions, evaluation scripts, and detailed evaluation reports)Extracted from: pdfAgreement 49%

Explore related topics

Related papers