12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Vector RAG vs LLM-Compiled Wiki: A Preregistered Comparison on a Small Multi-Domain Research

Theodore O. Cochran · ArXiv.org · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
1/4
Quality (LMQS)
E
Evidence
0
Citations
0.00
FWCI

Methodology & findings

Study design

Preregistered, blinded, two-judge comparison study.

Sample

N = 13, 8 groups

Primary method

Bootstrap percentile CI (10,000 resamples per criterion × tier); Bayesian posterior estimation with score_diff ∼ N(µ, σ²), µ ∼ N(0, 4), σ ∼ HalfNormal(4), NUTS sampling (4 chains × 2,000 iter × 2,000 warmup, random_seed=42). Convergence checked via R̂ ≤ 1.003 and min ESS ≥ 1,907. Strongly corroborated threshold: P(µ ≥ θ) ≥ 0.95. Inter-rater reliability (IRR) adjustment via judge-mean recomputation with preregistered max-delta trigger rule. Exploratory decomposition-retrieval ablation with three precommitted predictions (P1-P3). Post-hoc claim-level grounding analysis using binary/ternary classification (supported/partial/contradicted/unsupported).

Main result

The study found that "Wiki -RAG ≥ +2.0 on inter_paper_mapping" with an inter-paper mapping advantage that was "large and never close to the threshold," supporting the hypothesis that the LLM Wiki system provides better multi-hop synthesis. However, "the structural_integrity advantage is threshold-sensitive" (IRR-adjusted: +1.625, missing the +2.0 threshold by 0.375). Conversely, on point-source grounding (H2), RAG maintained parity with the wiki, but the claim-level analysis revealed that "wiki cited claims are ∼ 2× more often supported and ∼ 4-5× less often unsupported than RAG cited claims." Critically, "the wiki cost more per query, so the amortization story cannot hold," with a "21× per-query gap" in costs.

Reports effect sizes and confidence intervals.

Research paradigm

positivist empirical comparison

Author conclusions

"The most defensible conclusion of this preregistered comparison is not that compiled wiki memory beats RAG, nor the reverse. Grounded research synthesis is not a single capability: a system can organize evidence well, cite evidence well for each specific claim, or run cheaply, and in this experiment no architecture did all three best." The authors note that "Single-round RAG minimizes cost; decomp-RAG approaches the wiki on synthesis-shape rubric criteria at ∼ 3.4× lower per-query LLM-token cost; wiki retains the strongest LLM-scored evidence-artifact claim-citation alignment, at high per-query cost and pending source-PDF validation." They recommend that "Future small-corpus RAG/Wiki evaluations should report synthesis structure, claim-citation alignment, and cost separately rather than collapsing them into a single grounded-synthesis verdict."

Risk of bias

Judge calibration drift between GPT-5.4 and Gemini 2.5 Pro on holistic groundedness criterion; Gemini 2.5 Pro ceiling effects (ratings saturating near 10/10 on three of four criteria); LLM-as-judge biases documented in related literature (position bias, verbosity bias, self-enhancement bias); Small sample size (n=13 questions, n=4 for H1 confirmatory subset); Post-hoc analyses subject to multiple testing and judge sensitivity; Judge calibration drift: IRR analysis revealed max-deltas of 6 (groundedness), 4 (structural_integrity), 8 (conflict_awareness), and inter_paper_mapping differences, with Gemini 2.5 Pro ratings saturating near 10/10 and GPT-5.4 ratings spreading between 6-9; LLM-as-judge bias: consistent with prior literature documenting position, verbosity, and self-enhancement biases in LLM judges; ceiling-judge effect on holistic groundedness criterion; Architectural confounding: multiple differences between systems (offline LLM rewriting, markdown vs vector store, tool-loop vs single-round retrieval) prevent clean causal isolation; Single-round RAG limitation: the registered RAG configuration uses single-round retrieval, which is a known weakness on multi-hop synthesis per prior literature; Seed-based blinding potential drift: per-question blinding via random.Random(seed=42) and (seed=43) for secondary judge; Small sample size (n=13 questions, n=4 for confirmatory H1 subset): reduces statistical power and increases sensitivity to individual question outcomes; LLM-as-judge bias: documented saturation of Gemini 2.5 Pro ratings near 10/10 for wiki output on three criteria; Judge calibration drift: large disagreement on groundedness criterion (B2 question: GPT-5.4 scored RAG=9/wiki=6, Gemini 2.5 Pro scored RAG=4/wiki=10); Criterion definition sensitivity: holistic criteria (groundedness, structural_integrity) showed greater judge disagreement than concrete criteria (inter_paper_mapping); Query blinding design: Random blinding may not fully eliminate judge preference biases despite seed-based randomization; Post-hoc analysis bias: claim-level grounding analysis was not preregistered and reported as exploratory

Limitations

  • The study acknowledges that "for wiki claims, the cited evidence is a wiki-page excerpt, which is itself a compilation of source PDFs and not an original PDF passage," and "the analysis below therefore tests evidence-artifact claim alignment (does each claim follow from what the system cited?), not original-source fidelity (does each claim follow from the underlying PDF?)." Additionally, "the systems differ along multiple axes (offline LLM rewriting
  • compiled markdown vs vector store
  • tool-loop vs single-round retrieval
  • multi-call vs single-call context)," making it impossible to "a clean causal isolation of 'knowledge organization'." The study also notes that "even the wiki's cited claims are more often partial than strictly supported (53.1% vs 40.2% across all 13 questions)," and that "the two judges agreed on the sign of every criterion-level gap on the H1 confirmatory subset and on three of four 13-question criterion-level means," with groundedness showing the most judge calibration drift.

Open questions raised

  • Test source-PDF fidelity for wiki rather than only evidence-artifact alignment
  • Run Gemini 2.5 Pro on decomp-RAG pair to add IRR coverage
  • Test retrieval-during-generation methods that may close more of wiki advantage
  • Test hybrid wiki+source-RAG architectures for claim verification
  • Future small-corpus RAG/Wiki evaluations should report synthesis structure, claim-citation alignment, and cost separately rather than collapsing them into single verdict
  • Source-fidelity validation: "Future work should compare claim-level grounding against original PDF passages (testing source fidelity for wiki) rather than against the cited evidence artifact only"
Data: 24 peer-reviewed papers (2017-2026) across three domains: AI ethics & law, climate science, and precision medicine. Full corpus list with DOIs and evaluation questions available in OSF deposit.; 24 peer-reviewed papers across three domains (AI ethics & law, climate science, precision medicine, 2017-2026) with DOIs listed in OSF deposit; Corpus list with DOIs and full question text available in the OSF deposit; 466 atomic claims with claim-level grounding analysis available in OSF supplement; OSF deposit containing: corpus list with DOIs, full question text, all 13 evaluation questions, rubric criteria, supplementary prompts, and trial-level data (referenced but specific URL not extracted)Extracted from: pdfAgreement 54%

Explore related topics

Related papers