12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Causal-Enhanced AI Agents for Medical Research Screening

Duc Thinh Ngo, Arya Rahgoza · ArXiv.org · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
2/4
Quality (LMQS)
E
Evidence
0
Citations
0.00
FWCI

Methodology & findings

Study design

Comparative system evaluation using automated question-answering on a curated corpus of 234 peer-reviewed dementia exercise abstracts.

Sample

N = 234, 3 groups

Primary method

Manual human assessment using explicit rubrics on five dimensions: accuracy scores (0-10 ordinal scale), binary retrieval success determination, categorical evidence quality classification, hallucination detection (present/absent), and response characteristics documentation. No formal statistical inference testing or confidence intervals reported.

Main result

The study found that "CausalAgent achieved 95% accuracy, 100% retrieval success, and zero hallucinations versus 34% accuracy and 10% hallucinations for baseline AI." The evaluation on 234 dementia exercise abstracts demonstrates that "causal graph-enhanced retrieval-augmented generation fundamentally transforms AI-assisted medical evidence synthesis, achieving 95% accuracy with zero hallucinations through explicit causal reasoning integrated with retrieval mechanisms."

Reports effect sizes.

Research paradigm

Positivist/empiricist

Author conclusions

The authors conclude: "This research demonstrates that causal graph-enhanced retrieval-augmented generation fundamentally transforms AI-assisted medical evidence synthesis, achieving 95% accuracy with zero hallucinations through explicit causal reasoning integrated with retrieval mechanisms." They further state: "For health informatics, causal-enhanced RAG demonstrates how to augment rather than undermine clinical expertise. For AI research, it confirms that structured reasoning remains essential for high-stakes applications where errors have real-world consequences. The path forward is not black-box predictions replacing clinical judgment, but powerful tools making evidence more accessible, mechanisms more interpretable, and reasoning more transparent—a significant step toward human-AI collaboration in evidence-based healthcare."

Risk of bias

Single evaluator or limited inter-rater reliability in manual assessment; Domain-specific bias from exclusive focus on dementia exercise research; Question design bias - only 10 questions may not represent full scope of research QA scenarios; Temporal bias - results represent specific system versions at a specific time point; Small corpus size (234 abstracts) may not reflect performance on larger datasets; Single evaluator assessment - no inter-rater reliability reported; Single domain focus (dementia exercise only) - generalization bias; Question set bias - only 10 questions, may not represent full range of scenarios; Evaluator subjectivity in accuracy scoring despite rubric; Single evaluation instance per question - no variance assessment; Potential cherry-picking of questions favorable to RAG systems; Single evaluator assessment without inter-rater reliability metrics; Limited question set (10 questions) may not represent full spectrum of research QA scenarios; Single domain focus (dementia exercise research) limits generalizability; Single evaluation instance per question without temperature variations; Selection bias in corpus construction (PubMed and Google Scholar with specific filters)

Limitations

  • The study acknowledges several limitations: "Ten questions across four complexity levels may not fully represent all research QA scenarios
  • Future work should expand to all six complexity levels (including Level 5: Advanced Reasoning and Level 6: Meta-Analysis) with larger question sets (target: 200 questions as per original framework)." Additional limitations include: "Single Domain Focus: Evaluation used exclusively dementia exercise research abstracts
  • Generalization to other medical domains, scientific disciplines, or document types requires validation." "Manual Evaluation: Human assessment of accuracy and evidence quality introduces subjective judgment." "Corpus Size: 234 abstracts represents a focused corpus
  • Scalability to larger corpora (thousands or millions of documents) requires additional testing." "Binary Retrieval Metric: Success/failure classification does not capture partial retrieval quality, precision-recall tradeoffs, or retrieval ranking effectiveness." "Single Evaluation Instance: Each question was answered once per system
  • Multiple trials with temperature variations could reveal response consistency and variance."

Open questions raised

  • Expansion to all six complexity levels (200 questions total, including Level 5: Advanced Reasoning and Level 6: Meta-Analysis)
  • Multi-domain validation beyond dementia exercise research (oncology, cardiology mentioned)
  • Automated graph quality metrics development
  • Interactive evidence exploration capabilities
  • Real-time incremental updates for knowledge graphs
  • Scalability testing on larger corpora (thousands or millions of documents)
Data: Dementia exercise research corpus; Curated corpus of 234 dementia exercise abstracts from PubMed and Google Scholar (available on GitHub repository); Complete dataset and processing scripts available at: https://github.com/ngohuuduc/causalagents; Exercise interventions for dementia research corpusCode: GitHub: https://github.com/ngohuuduc/causalagents; https://github.com/ngohuuduc/causalagents (complete end-to-end implementation); CausalAgent implementationExtracted from: pdfAgreement 58%

Explore related topics

Related papers