MedResearchBench: A Multi-Domain Benchmark for Evaluating AI Research Agents on Clinical Medical Research
Shuping Tan, Zhanxiao Tian · medRxiv · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.64898/2026.03.30.26349749
Methodology & findings
Study design
Benchmark design and construction with pilot validation.
Sample
N = 3, 3 groups
Primary method
Weighted scoring methodology using LLM Judge evaluation. Objective mode: each checklist item scored 0-100 where 50 represents matching ground truth paper quality, with weighted sum across items producing final task score. Dimension-specific weights applied: statistical methodology (w=0.20), results accuracy (w=0.25), visualization quality (w=0.15), clinical interpretation (w=0.20), confounding sensitivity (w=0.10), reporting compliance (w=0.10). Supplementary subjective mode assessment of overall research quality, novelty, and clinical utility. Medical-specific extensions include survey design penalty and STROBE compliance checklist. Evaluation performed using GPT-4o or Claude Opus LLM Judge with prompt templates provided.
Main result
Initial baseline evaluation demonstrated that "End-to-end evaluation of an agentic pipeline across 3 pilot tasks (Tier 1-3), yielding a mean score of 72/100 (B-level), establishing the first quantitative baseline for AI-driven medical research quality on this benchmark." The pilot results showed survey-weighted methodology was correctly implemented with 100% compliance across all three tasks, but results accuracy was the primary limitation with a mean score of 58.7/100, driven by covariate incompleteness and reference group misspecification.
Reports effect sizes and confidence intervals.
Research paradigm
Positivist/empiricist - quantitative evaluation through systematic benchmarking
Author conclusions
The authors conclude: "As AI systems increasingly automate scientific research, the medical domain-with its high stakes, complex methodology, and established quality standards-demands specialized evaluation. MedResearchBench provides that evaluation platform and, through its anti-paper-mill design, serves as a quality gate for responsible AI-assisted medical research." They state the benchmark "is designed as a complement to ResearchClawBench, not a replacement. Together, the two benchmarks span the full landscape of AI research evaluation, from fundamental science to clinical medicine."
Risk of bias
Limited pilot evaluation (n=3 tasks only); Potential bias toward NHANES-based tasks (13/16 tasks); Ground truth paper selection bias: papers selected must meet specific quality criteria and use public datasets; LLM Judge evaluation dependency: scoring relies on GPT-4o or Claude Opus capabilities; Evaluation limited to analytical pipeline; does not capture full clinical research complexity; No inter-rater reliability reported for LLM Judge evaluation; Ground truth papers span wide IF range (2.3-51.0) but may not capture full distribution of real-world medical publications
Limitations
- The authors explicitly state: "No wet-lab or clinical data collection: Tasks evaluate the analytical pipeline only." Additional limitations are implied in the task design: "visualization quality appears in only 2 tasks in Phase 1
- Phase 2 will systematically increase visualization coverage." The benchmark is limited to Phase 1 tasks (16 tasks), and evaluation focuses on end-to-end analytical workflows from publicly available datasets rather than prospective data collection.
Open questions raised
- No existing benchmark tests whether an AI system can correctly go from patient-level data to a clinically sound, publication-ready manuscript
- Medical research tools remain comparatively underdeveloped compared to fundamental science automation
- No published system specifically targets end-to-end automation of observational clinical research
- Future work: Phase 2 will systematically increase visualization coverage (currently only 2/16 tasks)
- Potential for extending benchmark to prospective validation with new published papers
- No existing benchmark evaluates AI systems on medical clinical research tasks prior to MedResearchBench
Explore related topics
Related papers
- Which academic search systems are suitable for systematic reviews or meta‐analyses? Evaluating retrieval qualities of Google Scholar, PubMed, and 26 other resourcesMichael Gusenbauer · 2019 · 2,116 citations
- Artificial intelligence adoption in the physical sciences, natural sciences, life sciences, social sciences and the arts and humanities: A bibliometric analysis of research publications from 1960-2021Stefan Hajkowicz · 2023 · 119 citations
- PaperQA: Retrieval-Augmented Generative Agent for Scientific ResearchJakub Lála · 2023 · 52 citations
- Language agents achieve superhuman synthesis of scientific knowledgeMichael Skarlinski · 2024 · 40 citations
- Can ChatGPT be used to predict citation counts, readership, and social media interaction? An exploration among 2222 scientific abstractsJoost de Winter · 2024 · 29 citations
- Generative AI in Academic Writing: A Comparison of DeepSeek, Qwen, ChatGPT, Gemini, Llama, Mistral, and GemmaÖmer Aydın · 2025 · 17 citations