DeepER-Med: Advancing Deep Evidence-Based Research in Medicine Through Agentic AI
Zhizheng Wang, Chih-Hsuan Wei, Joey Chan, Robert Leaman, Chi-Ping Day, Chuan Wu et al. · arXiv (Cornell University) · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
Methodology & findings
Study design
Mixed-methods approach combining: (1) Expert-curated benchmark development with multidisciplinary panel (n=11 biomedical experts); (2) Comparative evaluation against three production-grade systems (OpenAI Deep Research, OpenEvidence, Google AI Mode/Deep Search) using blinded expert assessment across five evaluation dimensions; (3) Mechanistic analysis using five open-access biomedical datasets with quantitative metrics (semantic similarity, information entropy, Jensen-Shannon divergence); (4) Real-world clinical case study with three clinicians evaluating eight oncology cases from Precision Oncology Tumor Board..
Primary method
Design science research with participatory expert involvement; evidence-based design informed by medical research workflows and evidence-based medicine principles
Main result
DeepER-Med consistently outperformed production-grade platforms across multiple evaluation criteria. Specifically, "The largest improvements were observed in reference relevance (81 cases versus 59 cases for the strongest baseline) and analytical quality (67 cases versus 47 cases), indicating stronger evidence retrieval and integration." Additionally, "DeepER-Med was chosen in 60 cases, including 27 sole selections" when experts selected the best response among systems, compared to the strongest baseline system selected in 43 cases. In real-world clinical evaluation, "DeepER-Med's conclusions were judged consistent with previous human-expert recommendations of POTB in seven cases" out of eight evaluated cases.
Research paradigm
Design science research with empirical evaluation
Author conclusions
The authors conclude that "evidence-based generation (EBG) as a framework for structuring deep medical research with AI systems" represents a paradigm shift in AI-assisted biomedical research. They state: "These components shift the focus of AI-assisted biomedical research from task-level answer accuracy of regular LLM-only or RAG-based approaches toward the transparency and reliability of the deep research setting." They further conclude: "By decomposing research intent and enforcing explicit inclusion criteria, the system effectively 'opens the black box,' transforming scientific synthesis from an opaque heuristic into a transparent, auditable sequence of evidence-grounded decisions." Regarding clinical utility: "Rather than replacing multidisciplinary decision-making, such systems may support clinical discussion by identifying relevant studies and structuring evidence-intensive analyses."
Risk of bias
Expert selection bias: Questions contributed by participating experts may reflect their research interests; Evaluator bias: Experts evaluated responses to questions they formulated, though blinding was implemented; System selection bias: Only compared against three specific production-grade systems; other agentic systems not included; Publication bias in retrieved literature: System retrieves from existing databases, limited to published/indexed work; Clinical case study sample size: Only 8 cases evaluated, limiting generalizability; Expert selection bias: All 11 domain experts were from NIH, Johns Hopkins, or University of Illinois institutions, potentially limiting geographic and institutional diversity; Evaluator bias: While blinded annotation was employed, the expert who contributed each question conducted evaluation, potentially introducing subtle bias despite masking; Small sample size for case study: Only 8 clinical cases evaluated, limiting generalizability; Comparison system selection: Only three comparator systems evaluated; other agentic systems not included; Question formulation bias: Questions derived from participating experts' research focuses may not represent broader medical research landscape; Expert bias in question formulation and evaluation (though blinded annotation strategy was used to mitigate); Selection bias in clinical case study (only eight cases from single institution); Potential LLM-based judge bias (GPT-5.2-pro used in reference answer curation); Question author evaluated their own contributed questions (though blinded to system identity)
Limitations
- The paper acknowledges several limitations: "most deep research systems are evaluated using simplified questions drawn from open-access databases or extracted from scientific studies, which primarily emphasize multiple-choice answer accuracy without the verification of supporting references." Additionally, "the intermediate processes of evidence selection, aggregation, and interpretation are frequently opaque to researchers, making it difficult to determine whether the final conclusions reflect robust evidence synthesis." In the clinical evaluation, "The remaining case involved a clinically nuanced scenario with evolving evidence and was considered discordant," and "Evidence reliability was judged fully satisfactory in five cases and partially satisfactory in the remaining cases due to incomplete coverage of prior studies."
Open questions raised
- The authors identify several research gaps: (1) Lack of evaluation of deep research systems on complex, real-world medical questions; (2) Absence of explicit, inspectable criteria for evidence appraisal in existing systems; (3) Limited understanding of how AI-augmented research tools affect scientific productivity and diversity of explored topics; (4) Need for more rigorous benchmarking approaches that evaluate performance beyond simplified question-answer accuracy; (5) Insufficient examination of real-world performance in comprehensive literature exploration and expert-aligned evidence synthesis.
- Limited evaluation of real-world performance of deep research systems in comprehensive literature exploration and expert-aligned evidence synthesis
- Benchmarks primarily emphasize multiple-choice accuracy without verification of supporting references
- Lack of explicit and inspectable criteria for evidence appraisal in existing systems
- Need for evaluation capturing demands of frontline medical research in real-world scenarios with focus on trustworthiness and interpretability
- AI-augmented research tools may narrow diversity of topics collectively explored by researchers
Explore related topics
Related papers
- PRISMA Extension for Scoping Reviews (PRISMA-ScR): Checklist and ExplanationAndrea C. Tricco · 2018 · 40,391 citations
- Cochrane Handbook for Systematic Reviews of Interventions2019 · 14,420 citations
- Updated methodological guidance for the conduct of scoping reviewsMicah D.J. Peters · 2020 · 6,688 citations
- Role of AI chatbots in education: systematic literature reviewLasha Labadze · 2023 · 791 citations
- PaperQA: Retrieval-Augmented Generative Agent for Scientific ResearchJakub Lála · 2023 · 52 citations
- Co-designing AI Education Curriculum with Cross-Disciplinary High School TeachersBenjamin Xie · 2024 · 28 citations