12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

UniD³: A Knowledge Graph-Enhanced RAG Framework for Drug-Disease Discovery and Reasoning

Qing Wang, Tianshi Liu, Minghao Zhou, Jialu Liang, Sen Guo, Guangyu Wang et al. · arXiv (Cornell University) · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
3/4
Quality (LMQS)
D
Evidence
0
Citations
0.00
FWCI

Methodology & findings

Study design

Framework development combining Large Language Models (Llama 3.3-70B) with Knowledge Graph-enhanced Retrieval-Augmented Generation (KG-RAG).

Primary method

Design science research with iterative refinement through automated quality checks, fuzzy matching validation, and expert clinical review

Main result

UniD3 generates six knowledge graphs and large-scale datasets, including "28,915 DDM samples, 15,042 DEA samples, and over 4,000 DTA question–answer pairs." External validation demonstrates strong performance with "F1: 0.85–0.87 for DDM/DEA; 0.82 for DTA", and clinician review confirms high reliability with "AUROC = 0.90". Additionally, "KG-RAG–augmented models outperform standalone LLMs."

Research paradigm

Design science / computational informatics

Author conclusions

"UniD3 provides a scalable, extensible framework for transforming unstructured biomedical literature into high-quality, structured drug–disease knowledge, supporting AI-driven discovery, repurposing, and precision medicine." The authors further state that "UniD3 establishes a scalable, cost-effective, and extensible framework for transforming unstructured biomedical literature into structured Drug-Disease knowledge, supporting applications ranging from hypothesis generation to model benchmarking."

Risk of bias

Publication bias from PubMed dependence (positive/well-studied findings favored); LLM hallucination effects requiring fuzzy matching refinement; Context-specificity not generalizing across populations or disease subtypes; Potential for incomplete recognition and misclassification in automated extraction; Publication bias toward positive or well-studied findings in PubMed; LLM hallucination and misclassification in entity extraction; Incomplete entity recognition; Underrepresentation of rare conditions and under-reported therapeutic contexts; Context-specific associations that may not translate across populations or disease subtypes; LLM hallucination in entity and relationship extraction; Context-specific associations that may not translate across populations; Underrepresentation of rare diseases and under-reported therapeutic contexts

Limitations

  • "Because UniD3 relies on LLMs for entity and relationship extraction, issues such as misclassification, incomplete recognition, and occasional hallucination may still occur, potentially introducing noise into the knowledge graphs and datasets
  • Moreover, dependence on PubMed literature may introduce publication bias toward positive or well-studied findings, limiting coverage of rare conditions or under-reported therapeutic contexts
  • Although cross-dataset validation improves generalizability, certain associations remain context-specific and may not directly translate across populations or disease subtypes
  • At the knowledge-graph level, prioritization of direct Drug-Disease links may underrepresent multi-step biological pathways and indirect mechanistic relationships
  • Additionally, while KG-RAG grounding improves factual reliability, the LLM-driven extraction process still presents challenges for full transparency and explainability."

Open questions raised

  • Future work will focus on: (1) extending breadth and depth through continuous integration of newly published research and regulatory documents; (2) incorporating multi-modal biomedical data (molecular, genetic, clinical outcomes); (3) developing evidence-aware confidence scoring to distinguish high-confidence therapeutic associations from exploratory signals; (4) enhancing chatbot with multi-turn dialogue and user feedback mechanisms; (5) refining explainability mechanisms and expert-in-the-loop feedback for transparency and reliability.
  • Future work will focus on: (1) extending breadth and depth of knowledge graphs through continuous integration of newly published research and regulatory documents; (2) incorporating multi-modal biomedical data such as molecular, genetic, and clinical outcome evidence; (3) developing evidence-aware confidence scoring to distinguish high-confidence associations from exploratory signals; (4) enhancing chatbot with multi-turn dialogue capabilities and user feedback mechanisms; (5) further refining explainability mechanisms and expert-in-the-loop feedback.
  • Future work will focus on: (1) extending the breadth and depth of knowledge graphs through continuous integration of newly published research and regulatory documents; (2) incorporating multi-modal biomedical data such as molecular, genetic, and clinical outcome evidence; (3) developing evidence-aware confidence scoring to distinguish high-confidence therapeutic associations from exploratory research signals; (4) enhancing the chatbot with multi-turn dialogue capabilities and user feedback mechanisms; (5) further refinement of explainability mechanisms and expert-in-the-loop feedback.
Data: Drug-Disease Matching (DDM) dataset: https://huggingface.co/datasets/Mike2481/UniD3_DDM; Drug-Target Analysis (DTA) dataset: https://huggingface.co/datasets/Mike2481/UniD3_DTA; Drug Effectiveness Assessment (DEA) dataset: https://huggingface.co/datasets/Mike2481/UniD3_DEA; External benchmark datasets: CDR, DrugReviews, P3ps, MedicationQA (from Hugging Face); Drug-Disease Matching (DDM) dataset at https://huggingface.co/datasets/Mike2481/UniD3_DDM; Drug-Target Analysis (DTA) dataset at https://huggingface.co/datasets/Mike2481/UniD3_DTA; Drug Effectiveness Assessment (DEA) dataset at https://huggingface.co/datasets/Mike2481/UniD3_DEA; External benchmark datasets: CDR, DrugReviews, P3ps, MedicationQA from Hugging FaceCode: https://github.com/QSong-github/UniD3 (MIT License)Extracted from: pdfAgreement 61%

Explore related topics

Related papers