UniD³: A Knowledge Graph-Enhanced RAG Framework for Drug-Disease Discovery and Reasoning
Qing Wang, Tianshi Liu, Minghao Zhou, Jialu Liang, Sen Guo, Guangyu Wang et al. · arXiv (Cornell University) · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
Methodology & findings
Study design
Framework development combining Large Language Models (Llama 3.3-70B) with Knowledge Graph-enhanced Retrieval-Augmented Generation (KG-RAG).
Primary method
Design science research with iterative refinement through automated quality checks, fuzzy matching validation, and expert clinical review
Main result
UniD3 generates six knowledge graphs and large-scale datasets, including "28,915 DDM samples, 15,042 DEA samples, and over 4,000 DTA question–answer pairs." External validation demonstrates strong performance with "F1: 0.85–0.87 for DDM/DEA; 0.82 for DTA", and clinician review confirms high reliability with "AUROC = 0.90". Additionally, "KG-RAG–augmented models outperform standalone LLMs."
Research paradigm
Design science / computational informatics
Author conclusions
"UniD3 provides a scalable, extensible framework for transforming unstructured biomedical literature into high-quality, structured drug–disease knowledge, supporting AI-driven discovery, repurposing, and precision medicine." The authors further state that "UniD3 establishes a scalable, cost-effective, and extensible framework for transforming unstructured biomedical literature into structured Drug-Disease knowledge, supporting applications ranging from hypothesis generation to model benchmarking."
Risk of bias
Publication bias from PubMed dependence (positive/well-studied findings favored); LLM hallucination effects requiring fuzzy matching refinement; Context-specificity not generalizing across populations or disease subtypes; Potential for incomplete recognition and misclassification in automated extraction; Publication bias toward positive or well-studied findings in PubMed; LLM hallucination and misclassification in entity extraction; Incomplete entity recognition; Underrepresentation of rare conditions and under-reported therapeutic contexts; Context-specific associations that may not translate across populations or disease subtypes; LLM hallucination in entity and relationship extraction; Context-specific associations that may not translate across populations; Underrepresentation of rare diseases and under-reported therapeutic contexts
Limitations
- "Because UniD3 relies on LLMs for entity and relationship extraction, issues such as misclassification, incomplete recognition, and occasional hallucination may still occur, potentially introducing noise into the knowledge graphs and datasets
- Moreover, dependence on PubMed literature may introduce publication bias toward positive or well-studied findings, limiting coverage of rare conditions or under-reported therapeutic contexts
- Although cross-dataset validation improves generalizability, certain associations remain context-specific and may not directly translate across populations or disease subtypes
- At the knowledge-graph level, prioritization of direct Drug-Disease links may underrepresent multi-step biological pathways and indirect mechanistic relationships
- Additionally, while KG-RAG grounding improves factual reliability, the LLM-driven extraction process still presents challenges for full transparency and explainability."
Open questions raised
- Future work will focus on: (1) extending breadth and depth through continuous integration of newly published research and regulatory documents; (2) incorporating multi-modal biomedical data (molecular, genetic, clinical outcomes); (3) developing evidence-aware confidence scoring to distinguish high-confidence therapeutic associations from exploratory signals; (4) enhancing chatbot with multi-turn dialogue and user feedback mechanisms; (5) refining explainability mechanisms and expert-in-the-loop feedback for transparency and reliability.
- Future work will focus on: (1) extending breadth and depth of knowledge graphs through continuous integration of newly published research and regulatory documents; (2) incorporating multi-modal biomedical data such as molecular, genetic, and clinical outcome evidence; (3) developing evidence-aware confidence scoring to distinguish high-confidence associations from exploratory signals; (4) enhancing chatbot with multi-turn dialogue capabilities and user feedback mechanisms; (5) further refining explainability mechanisms and expert-in-the-loop feedback.
- Future work will focus on: (1) extending the breadth and depth of knowledge graphs through continuous integration of newly published research and regulatory documents; (2) incorporating multi-modal biomedical data such as molecular, genetic, and clinical outcome evidence; (3) developing evidence-aware confidence scoring to distinguish high-confidence therapeutic associations from exploratory research signals; (4) enhancing the chatbot with multi-turn dialogue capabilities and user feedback mechanisms; (5) further refinement of explainability mechanisms and expert-in-the-loop feedback.
Explore related topics
Related papers
- Rayyan—a web and mobile app for systematic reviewsMourad Ouzzani · 2016 · 24,664 citations
- Guidance for conducting systematic scoping reviewsMicah D.J. Peters · 2015 · 7,472 citations
- Which academic search systems are suitable for systematic reviews or meta‐analyses? Evaluating retrieval qualities of Google Scholar, PubMed, and 26 other resourcesMichael Gusenbauer · 2019 · 2,116 citations
- Hallucination Rates and Reference Accuracy of ChatGPT and Bard for Systematic Reviews: Comparative AnalysisMikaël Chelli · 2024 · 295 citations
- Artificial intelligence and the conduct of literature reviewsGerit Wagner · 2021 · 275 citations
- Artificial intelligence for literature reviews: opportunities and challengesF. J. Bolaños · 2024 · 188 citations