12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

VaaS is a Multi-Layer Hallucination Reduction Pipeline for AI-Assisted Science: Production Validation and Prospective Benchmarking

Ankit Sabharwal, Milit S Patel, Anna Carrano, Maarten Rotman, Wesley A. Wierson, Stephen C. Ekker · medRxiv · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
2/4
Quality (LMQS)
E
Evidence
0
Citations
0.00
FWCI

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.64898/2026.03.24.26348935

Methodology & findings

Study design

Multi-method empirical validation including: (1) production deployment across three waves (225 annotated rare disease gene curations and prospective 100-gene collection with over 3,000 verified citations); (2) controlled stress test battery (VaaS-HT-001) with five genes under unguided and protocol conditions; (3) prospective ablation study (VaaS-RIKER2) with 640 controlled runs (4 conditions × 4 temperatures × 40 genes) plus 117 open-weight model runs on dedicated GPU hardware; (4) independent L3 citation audit of Wave 3 (179 PMIDs, 100 genes); (5) evaluation against MedHallu clinical hallucination benchmark with three independent runs..

Sample

N = 757, 23 groups

Primary method

Hypothesis testing: Fisher's exact test comparing unguided vs. VaaS protocol (p < 0.001). Confidence interval calculation: 95% Wilson confidence intervals reported for error rate proportions. Levenshtein similarity scoring: ≥0.85 threshold for title matching in live fetch verification. Cross-model comparison: descriptive analysis across architectures (no formal statistical tests across models reported). Power analysis: indicated n=30 minimum for detecting reduction from 20% to 0% error rate with 80% power at α=0.05 (stress test was underpowered at n=5).

Main result

The study found that "the net functional hallucination rate approached zero" after three iterations of directed refinement of the VaaS pipeline. More specifically, "the full VaaS protocol achieved 0.0% Type I and 6.5% Type II, a >14-fold reduction" compared to unguided output which "produced 95.9% Type II hallucination." In the prospective VaaS-RIKER2 benchmark with 757 total runs (640 Claude + 117 open-weight), "the full VaaS pipeline (C4) reduced the wrong-topic rate to 6.5% (intercepted and rejected by the verification gate, not errors in the final output)," and on the MedHallu benchmark "the VaaS protocol achieved F1 = 0.9853 on the hard tier."

Reports effect sizes and confidence intervals.

Research paradigm

Empirical-computational (hybrid design with AI agents and human expert validation)

Author conclusions

"A rigorous, multi-layer AI science quality pipeline can achieve near-zero citation fabrication rates and substantially reduce drug approval hallucinations at production scale." The authors conclude that "The VaaS-RIKER2 prospective benchmark provides the first controlled ablation study of such a pipeline in rare disease AI synthesis (757 total runs: 640 Claude arm + 117 open-weight model arm), confirming three of four prospectively defined hypotheses." They establish "Structural universality: Wrong-topic citation hallucination is model-agnostic" and "Verification is the load-bearing gate: C3 (live PMID verification only) achieves 0.0%/0.0% by catching wrong-topic and non-existent candidates before output." Finally, "The VaaS approach represents, to the authors' knowledge, the lowest measured hallucination system for science to date and is set to further accelerate the use of AI and AI agents for advancing research."

Risk of bias

Selection bias: Stress test genes (VaaS-HT-001) were selected post-hoc from production database, not held-out prospectively; Self-evaluation bias: In MedHallu benchmark, models evaluated hallucinations generated by same model family; Ordering bias: Ground-truth answer consistently presented as option A in all MedHallu forced-choice comparisons; Lack of blinding: Human validation reviews not conducted blind to pipeline condition; Scoring tool anomaly: Open-weight arm (CHCHD10 runs) produced spurious valid classifications affecting valid rate calculations; C2 measurement methodology weakness: Type I assessment via AI self-reported calibration rather than live PubMed verification; No inter-rater reliability formally measured for human validation; Post-hoc gene selection in stress test (VaaS-HT-001) - genes drawn from production database rather than held-out set; No blinding in human validation reviews; Inter-rater reliability (IRR) not formally measured; Self-evaluation bias in MedHallu evaluation - model evaluated hallucinations generated by same model family; Ordering bias in MedHallu evaluation - ground-truth presented as option A in all batches; Scoring tool anomaly affecting open-weight model arm (CHCHD10 runs); C2 Type I measurement using AI self-reported calibration rather than live verification, not directly comparable to other conditions; Cross-study comparison issues - MedHallu evaluation uses forced-choice format while published baselines use per-answer binary classification; Selection bias: Production audit used post-hoc gene selection from same database pipeline produced, not held-out test set; Self-evaluation bias: MedHallu Run 2 evaluation had model evaluating hallucinations from same model family; Ordering bias: MedHallu evaluations with ground-truth answer presented consistently as option A; Scoring tool anomaly: CHCHD10 open-weight runs affected by spurious validity classifications; Single-operator replication: MedHallu Run 3 conducted by single independent operator; No inter-rater reliability: Human validation not conducted blind to pipeline condition; C2 measurement methodology: Type I assessment via AI self-reported calibration rather than live verification; Cost figure uncertainty: API costs derived from aggregate billing records with parallel projects

Limitations

  • "The VaaS-HT-001 stress test used n = 5 genes, which is substantially underpowered for formal hypothesis testing (minimum n = 30 indicated by power analysis)
  • Results should be interpreted as directional." Additionally, "The RIKER2 gene manifest was drawn exclusively from the mitochondrial disease core of the database
  • Generalizability to other disease categories (lysosomal, cardiac, renal, sensory) is inferred from Wave 1/2 production data but not formally benchmarked." The authors note that "The pipeline intercepts citation fabrication and wrong-topic citations
  • It does not intercept interpretation errors - cases where a real, on-topic paper is cited but the stated claim misrepresents the paper's finding
  • This error category requires domain expert review." Furthermore, "The reported ~$1/gene figure reflects API token costs only
  • Compute, infrastructure, and human review costs are excluded

Open questions raised

  • The authors identify the following gaps and future directions: (1) generalizability beyond rare disease genetics to other biomedical domains - "Generalizability to other biomedical domains has not been tested and cannot be assumed"; (2) interpretation error detection - "The pipeline intercepts citation fabrication and wrong-topic citations. It does not intercept interpretation errors"; (3) need for blinded evaluation - "Future work should incorporate blinded evaluation and IRR scoring for ambiguous (BORDERLINE) citations"; (4) domain-specific adaptation requirements - "adaptation to other domains requires domain-specific revalidation"; (5) scalability of human expert review loop - "The self-improving loop is only as good as the human catches that seed it. Domain expertise is not optional in the VaaS architecture."
  • The authors identify that "Zero hallucinations remain a theoretical limit, not a practical target." They note that interpretation error detection "requires domain expertise that no automated fetch can provide" and that the pipeline requires "human scientists contribute: (1) identification of novel error patterns; (2) verification of borderline citation decisions; (3) interpretation of multi-study findings; and (4) final sign-off on database entries." Future work is recommended to "incorporate blinded evaluation and IRR scoring for ambiguous (BORDERLINE) citations." Generalizability questions remain: "Generalizability to other disease categories (lysosomal, cardiac, renal, sensory) is inferred from Wave 1/2 production data but not formally benchmarked" and "Generalizability to other biomedical domains has not been tested and cannot be assumed."
  • Authors identify need for: (1) domain-specific adaptation of topic verification prompts and tiered risk criteria for non-rare-disease biomedical domains; (2) interpretation error detection requiring domain expertise beyond automated verification; (3) generalization testing across disease categories beyond mitochondrial/neurological/lysosomal/cardiac/sensory/immune/renal; (4) blinded evaluation and formal inter-rater reliability measurement in future work; (5) test-retest reliability measurement for MedHallu evaluations; (6) investigation of CDER/CBER FDA jurisdictional split as systematic knowledge gap.
Data: Rare Disease Database (RDD) v1.0 - described as 'open-source, living' database with 225 annotated rare disease gene curations plus prospective 100-gene collection; VaaS-RIKER2 40-gene manifest with verified PMID sets (5 PMIDs per gene; n=200 ground-truth PMIDs) curated in manifest.json; MedHallu benchmark (10,000 PubMedQA-derived question-answer pairs) - external benchmark from Pandit et al., 2025; The paper mentions "an open-source, living Rare Disease Database (RDD)" and references "a 40-gene manifest constructed from the mitochondrial disease core of the v1.0 database" with "verified PMID sets (5 PMIDs per gene; n = 200 total ground-truth PMIDs) curated in manifest.json," but specific URLs or accession numbers for public access are not provided in the text.; Rare Disease Database (RDD) v1.0 mentioned as open-source living database; specific public URL/DOI not provided in text. 40-gene manifest for VaaS-RIKER2 stored in manifest.json (format mentioned but repository location not provided). MedHallu benchmark data sourced from Pandit et al. (2025); original PubMedQA-derived.Code: Not explicitly mentioned in document; Not explicitly mentioned in the provided text.; No GitHub, GitLab, or other code repositories explicitly mentioned with URLsExtracted from: pdfAgreement 41%

Explore related topics

Related papers