MedResearchBench: A Multi-Domain Benchmark for Evaluating AI Research Agents on Clinical Medical Research
Shuping Tan, Zhanxiao Tian · medRxiv · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.64898/2026.03.30.26349749
Methodology & findings
Study design
Benchmark design and construction with pilot validation.
Sample
N = 3, 7 groups
Primary method
Weighted scoring methodology using LLM Judge evaluation. Objective mode: each checklist item scored 0-100 where 50 represents matching ground truth paper quality, with weighted sum across items producing final task score. Dimension-specific weights applied: statistical methodology (w=0.20), results accuracy (w=0.25), visualization quality (w=0.15), clinical interpretation (w=0.20), confounding sensitivity (w=0.10), reporting compliance (w=0.10). Supplementary subjective mode assessment of overall research quality, novelty, and clinical utility. Medical-specific extensions include survey design penalty and STROBE compliance checklist. Evaluation performed using GPT-4o or Claude Opus LLM Judge with prompt templates provided.
Main result
Initial baseline evaluation demonstrated that "End-to-end evaluation of an agentic pipeline across 3 pilot tasks (Tier 1-3), yielding a mean score of 72/100 (B-level), establishing the first quantitative baseline for AI-driven medical research quality on this benchmark." The pilot results showed survey-weighted methodology was correctly implemented with 100% compliance across all three tasks, but results accuracy was the primary limitation with a mean score of 58.7/100, driven by covariate incompleteness and reference group misspecification.
Reports effect sizes and confidence intervals.
Research paradigm
Positivist/empiricist - quantitative evaluation through systematic benchmarking
Author conclusions
The authors conclude: "As AI systems increasingly automate scientific research, the medical domain-with its high stakes, complex methodology, and established quality standards-demands specialized evaluation. MedResearchBench provides that evaluation platform and, through its anti-paper-mill design, serves as a quality gate for responsible AI-assisted medical research." They state the benchmark "is designed as a complement to ResearchClawBench, not a replacement. Together, the two benchmarks span the full landscape of AI research evaluation, from fundamental science to clinical medicine."
Risk of bias
Limited pilot evaluation (n=3 tasks only); Potential bias toward NHANES-based tasks (13/16 tasks); Ground truth paper selection bias: papers selected must meet specific quality criteria and use public datasets; LLM Judge evaluation dependency: scoring relies on GPT-4o or Claude Opus capabilities; Evaluation limited to analytical pipeline; does not capture full clinical research complexity; Selection bias in ground truth paper selection (papers must meet authenticity, quality, relevance, methodological rigor criteria and anti-paper-mill filtering). Task difficulty stratification may introduce complexity bias. LLM Judge evaluation introduces potential bias depending on model choice (GPT-4o or Claude Opus). Evaluation limited to publicly available datasets may not represent all clinical research paradigms.; Selection bias in ground truth paper selection (papers meeting specific criteria may not represent full diversity of medical research quality); LLM Judge bias (evaluation depends on GPT-4o or Claude Opus behavior, which may have inherent biases); Limited baseline validation (only 3 tasks evaluated, not the full 16-task benchmark); NHANES dominance (13/16 tasks) may introduce dataset-specific bias; No inter-rater reliability reported for LLM Judge evaluation; Ground truth papers span wide IF range (2.3-51.0) but may not capture full distribution of real-world medical publications
Limitations
- The authors explicitly state: "No wet-lab or clinical data collection: Tasks evaluate the analytical pipeline only." Additional limitations are implied in the task design: "visualization quality appears in only 2 tasks in Phase 1
- Phase 2 will systematically increase visualization coverage." The benchmark is limited to Phase 1 tasks (16 tasks), and evaluation focuses on end-to-end analytical workflows from publicly available datasets rather than prospective data collection.
Open questions raised
- No existing benchmark tests whether an AI system can correctly go from patient-level data to a clinically sound, publication-ready manuscript
- Medical research tools remain comparatively underdeveloped compared to fundamental science automation
- No published system specifically targets end-to-end automation of observational clinical research
- Future work: Phase 2 will systematically increase visualization coverage (currently only 2/16 tasks)
- Potential for extending benchmark to prospective validation with new published papers
- The authors identify that no existing benchmark tests "whether an AI system can correctly go from patient-level data to a clinically sound, publication-ready manuscript." They note that ResearchClawBench, the most comprehensive existing benchmark, "excludes clinical medicine entirely" from its 10 fundamental science domains. Medical research tools remain "comparatively underdeveloped" with "no published system specifically target[ing] the end-to-end automation of observational clinical research." Future directions mentioned include systematic increase of visualization quality coverage (Phase 2 will increase from 2 to more tasks), and extension beyond Phase 1's 16 tasks.
Explore related topics
Related papers
- Which academic search systems are suitable for systematic reviews or meta‐analyses? Evaluating retrieval qualities of Google Scholar, PubMed, and 26 other resourcesMichael Gusenbauer · 2019 · 2,116 citations
- Artificial intelligence adoption in the physical sciences, natural sciences, life sciences, social sciences and the arts and humanities: A bibliometric analysis of research publications from 1960-2021Stefan Hajkowicz · 2023 · 119 citations
- PaperQA: Retrieval-Augmented Generative Agent for Scientific ResearchJakub Lála · 2023 · 52 citations
- Language agents achieve superhuman synthesis of scientific knowledgeMichael Skarlinski · 2024 · 40 citations
- Can ChatGPT be used to predict citation counts, readership, and social media interaction? An exploration among 2222 scientific abstractsJoost de Winter · 2024 · 29 citations
- Generative AI in Academic Writing: A Comparison of DeepSeek, Qwen, ChatGPT, Gemini, Llama, Mistral, and GemmaÖmer Aydın · 2025 · 17 citations