12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

DeepER-Med: Advancing Deep Evidence-Based Research in Medicine Through Agentic AI

Zhizheng Wang, Chih-Hsuan Wei, Joey Chan, Robert Leaman, Chi-Ping Day, Chuan Wu et al. · arXiv (Cornell University) · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
1/4
Quality (LMQS)
D
Evidence
0
Citations
0.00
FWCI

Methodology & findings

Study design

Mixed-methods approach combining: (1) Expert-curated benchmark development with multidisciplinary panel (n=11 biomedical experts); (2) Comparative evaluation against three production-grade systems (OpenAI Deep Research, OpenEvidence, Google AI Mode/Deep Search) using blinded expert assessment across five evaluation dimensions; (3) Mechanistic analysis using five open-access biomedical datasets with quantitative metrics (semantic similarity, information entropy, Jensen-Shannon divergence); (4) Real-world clinical case study with three clinicians evaluating eight oncology cases from Precision Oncology Tumor Board..

Primary method

Design science research with participatory expert involvement; evidence-based design informed by medical research workflows and evidence-based medicine principles

Main result

DeepER-Med consistently outperformed production-grade platforms across multiple evaluation criteria. Specifically, "The largest improvements were observed in reference relevance (81 cases versus 59 cases for the strongest baseline) and analytical quality (67 cases versus 47 cases), indicating stronger evidence retrieval and integration." Additionally, "DeepER-Med was chosen in 60 cases, including 27 sole selections" when experts selected the best response among systems, compared to the strongest baseline system selected in 43 cases. In real-world clinical evaluation, "DeepER-Med's conclusions were judged consistent with previous human-expert recommendations of POTB in seven cases" out of eight evaluated cases.

Research paradigm

Design science research with empirical evaluation

Author conclusions

The authors conclude that "evidence-based generation (EBG) as a framework for structuring deep medical research with AI systems" represents a paradigm shift in AI-assisted biomedical research. They state: "These components shift the focus of AI-assisted biomedical research from task-level answer accuracy of regular LLM-only or RAG-based approaches toward the transparency and reliability of the deep research setting." They further conclude: "By decomposing research intent and enforcing explicit inclusion criteria, the system effectively 'opens the black box,' transforming scientific synthesis from an opaque heuristic into a transparent, auditable sequence of evidence-grounded decisions." Regarding clinical utility: "Rather than replacing multidisciplinary decision-making, such systems may support clinical discussion by identifying relevant studies and structuring evidence-intensive analyses."

Risk of bias

Expert selection bias: Questions contributed by participating experts may reflect their research interests; Evaluator bias: Experts evaluated responses to questions they formulated, though blinding was implemented; System selection bias: Only compared against three specific production-grade systems; other agentic systems not included; Publication bias in retrieved literature: System retrieves from existing databases, limited to published/indexed work; Clinical case study sample size: Only 8 cases evaluated, limiting generalizability; Expert selection bias: All 11 domain experts were from NIH, Johns Hopkins, or University of Illinois institutions, potentially limiting geographic and institutional diversity; Evaluator bias: While blinded annotation was employed, the expert who contributed each question conducted evaluation, potentially introducing subtle bias despite masking; Small sample size for case study: Only 8 clinical cases evaluated, limiting generalizability; Comparison system selection: Only three comparator systems evaluated; other agentic systems not included; Question formulation bias: Questions derived from participating experts' research focuses may not represent broader medical research landscape; Expert bias in question formulation and evaluation (though blinded annotation strategy was used to mitigate); Selection bias in clinical case study (only eight cases from single institution); Potential LLM-based judge bias (GPT-5.2-pro used in reference answer curation); Question author evaluated their own contributed questions (though blinded to system identity)

Limitations

  • The paper acknowledges several limitations: "most deep research systems are evaluated using simplified questions drawn from open-access databases or extracted from scientific studies, which primarily emphasize multiple-choice answer accuracy without the verification of supporting references." Additionally, "the intermediate processes of evidence selection, aggregation, and interpretation are frequently opaque to researchers, making it difficult to determine whether the final conclusions reflect robust evidence synthesis." In the clinical evaluation, "The remaining case involved a clinically nuanced scenario with evolving evidence and was considered discordant," and "Evidence reliability was judged fully satisfactory in five cases and partially satisfactory in the remaining cases due to incomplete coverage of prior studies."

Open questions raised

  • The authors identify several research gaps: (1) Lack of evaluation of deep research systems on complex, real-world medical questions; (2) Absence of explicit, inspectable criteria for evidence appraisal in existing systems; (3) Limited understanding of how AI-augmented research tools affect scientific productivity and diversity of explored topics; (4) Need for more rigorous benchmarking approaches that evaluate performance beyond simplified question-answer accuracy; (5) Insufficient examination of real-world performance in comprehensive literature exploration and expert-aligned evidence synthesis.
  • Limited evaluation of real-world performance of deep research systems in comprehensive literature exploration and expert-aligned evidence synthesis
  • Benchmarks primarily emphasize multiple-choice accuracy without verification of supporting references
  • Lack of explicit and inspectable criteria for evidence appraisal in existing systems
  • Need for evaluation capturing demands of frontline medical research in real-world scenarios with focus on trustworthiness and interpretability
  • AI-augmented research tools may narrow diversity of topics collectively explored by researchers
Data: DeepER-MedQA: 100 expert-curated medical research questions (availability not explicitly stated in paper); PubMedQA: 500 questions (referenced as external dataset); BioMaze: 2,523 open-ended QA questions (referenced as external dataset); MedAESQA: 40 evidence interpretation questions (referenced as external dataset); BioDSA: 520 evidence interpretation questions (referenced as external dataset); BioMaze Binary QA: 2,622 binary choice questions (referenced as external dataset); DeepER-MedQA: 100 expert-curated benchmark questions (availability not explicitly stated); PubMedQA: 500 questions with context; BioMaze Open QA: 2,523 open-ended biomedical QA questions; MedAESQA: 40 attribution QA questions; BioDSA: 520 hypothesis verification questions; BioMaze Binary QA: 2,622 binary choice questions; DeepER-MedQA: 100 expert-curated questions (availability not explicitly stated with URL); PubMedQA: 500 questions (referenced as publicly available dataset); BioMaze: 2,523 questions for open-ended QA (publicly available); MedAESQA: 40 questions (publicly available); BioDSA: 520 questions subset (publicly available); BioMaze Binary QA: 2,622 questions (publicly available); Clinical trial data from ClinicalTrials.gov; Literature from PubMedCode: No code repositories explicitly mentioned in the paper; Not mentioned in the paper excerpt provided.Extracted from: pdfAgreement 44%

Explore related topics

Related papers