Vector RAG vs LLM-Compiled Wiki: A Preregistered Comparison on a Small Multi-Domain Research
Theodore O. Cochran · ArXiv.org · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
Methodology & findings
Study design
Preregistered, blinded, two-judge comparison study.
Sample
N = 13, 8 groups
Primary method
Bootstrap percentile CI (10,000 resamples per criterion × tier); Bayesian posterior estimation with score_diff ∼ N(µ, σ²), µ ∼ N(0, 4), σ ∼ HalfNormal(4), NUTS sampling (4 chains × 2,000 iter × 2,000 warmup, random_seed=42). Convergence checked via R̂ ≤ 1.003 and min ESS ≥ 1,907. Strongly corroborated threshold: P(µ ≥ θ) ≥ 0.95. Inter-rater reliability (IRR) adjustment via judge-mean recomputation with preregistered max-delta trigger rule. Exploratory decomposition-retrieval ablation with three precommitted predictions (P1-P3). Post-hoc claim-level grounding analysis using binary/ternary classification (supported/partial/contradicted/unsupported).
Main result
The study found that "Wiki -RAG ≥ +2.0 on inter_paper_mapping" with an inter-paper mapping advantage that was "large and never close to the threshold," supporting the hypothesis that the LLM Wiki system provides better multi-hop synthesis. However, "the structural_integrity advantage is threshold-sensitive" (IRR-adjusted: +1.625, missing the +2.0 threshold by 0.375). Conversely, on point-source grounding (H2), RAG maintained parity with the wiki, but the claim-level analysis revealed that "wiki cited claims are ∼ 2× more often supported and ∼ 4-5× less often unsupported than RAG cited claims." Critically, "the wiki cost more per query, so the amortization story cannot hold," with a "21× per-query gap" in costs.
Reports effect sizes and confidence intervals.
Research paradigm
positivist empirical comparison
Author conclusions
"The most defensible conclusion of this preregistered comparison is not that compiled wiki memory beats RAG, nor the reverse. Grounded research synthesis is not a single capability: a system can organize evidence well, cite evidence well for each specific claim, or run cheaply, and in this experiment no architecture did all three best." The authors note that "Single-round RAG minimizes cost; decomp-RAG approaches the wiki on synthesis-shape rubric criteria at ∼ 3.4× lower per-query LLM-token cost; wiki retains the strongest LLM-scored evidence-artifact claim-citation alignment, at high per-query cost and pending source-PDF validation." They recommend that "Future small-corpus RAG/Wiki evaluations should report synthesis structure, claim-citation alignment, and cost separately rather than collapsing them into a single grounded-synthesis verdict."
Risk of bias
Judge calibration drift between GPT-5.4 and Gemini 2.5 Pro on holistic groundedness criterion; Gemini 2.5 Pro ceiling effects (ratings saturating near 10/10 on three of four criteria); LLM-as-judge biases documented in related literature (position bias, verbosity bias, self-enhancement bias); Small sample size (n=13 questions, n=4 for H1 confirmatory subset); Post-hoc analyses subject to multiple testing and judge sensitivity; Judge calibration drift: IRR analysis revealed max-deltas of 6 (groundedness), 4 (structural_integrity), 8 (conflict_awareness), and inter_paper_mapping differences, with Gemini 2.5 Pro ratings saturating near 10/10 and GPT-5.4 ratings spreading between 6-9; LLM-as-judge bias: consistent with prior literature documenting position, verbosity, and self-enhancement biases in LLM judges; ceiling-judge effect on holistic groundedness criterion; Architectural confounding: multiple differences between systems (offline LLM rewriting, markdown vs vector store, tool-loop vs single-round retrieval) prevent clean causal isolation; Single-round RAG limitation: the registered RAG configuration uses single-round retrieval, which is a known weakness on multi-hop synthesis per prior literature; Seed-based blinding potential drift: per-question blinding via random.Random(seed=42) and (seed=43) for secondary judge; Small sample size (n=13 questions, n=4 for confirmatory H1 subset): reduces statistical power and increases sensitivity to individual question outcomes; LLM-as-judge bias: documented saturation of Gemini 2.5 Pro ratings near 10/10 for wiki output on three criteria; Judge calibration drift: large disagreement on groundedness criterion (B2 question: GPT-5.4 scored RAG=9/wiki=6, Gemini 2.5 Pro scored RAG=4/wiki=10); Criterion definition sensitivity: holistic criteria (groundedness, structural_integrity) showed greater judge disagreement than concrete criteria (inter_paper_mapping); Query blinding design: Random blinding may not fully eliminate judge preference biases despite seed-based randomization; Post-hoc analysis bias: claim-level grounding analysis was not preregistered and reported as exploratory
Limitations
- The study acknowledges that "for wiki claims, the cited evidence is a wiki-page excerpt, which is itself a compilation of source PDFs and not an original PDF passage," and "the analysis below therefore tests evidence-artifact claim alignment (does each claim follow from what the system cited?), not original-source fidelity (does each claim follow from the underlying PDF?)." Additionally, "the systems differ along multiple axes (offline LLM rewriting
- compiled markdown vs vector store
- tool-loop vs single-round retrieval
- multi-call vs single-call context)," making it impossible to "a clean causal isolation of 'knowledge organization'." The study also notes that "even the wiki's cited claims are more often partial than strictly supported (53.1% vs 40.2% across all 13 questions)," and that "the two judges agreed on the sign of every criterion-level gap on the H1 confirmatory subset and on three of four 13-question criterion-level means," with groundedness showing the most judge calibration drift.
Open questions raised
- Test source-PDF fidelity for wiki rather than only evidence-artifact alignment
- Run Gemini 2.5 Pro on decomp-RAG pair to add IRR coverage
- Test retrieval-during-generation methods that may close more of wiki advantage
- Test hybrid wiki+source-RAG architectures for claim verification
- Future small-corpus RAG/Wiki evaluations should report synthesis structure, claim-citation alignment, and cost separately rather than collapsing them into single verdict
- Source-fidelity validation: "Future work should compare claim-level grounding against original PDF passages (testing source fidelity for wiki) rather than against the cited evidence artifact only"
Explore related topics
Related papers
- PRISMA Extension for Scoping Reviews (PRISMA-ScR): Checklist and ExplanationAndrea C. Tricco · 2018 · 40,391 citations
- Cochrane Handbook for Systematic Reviews of Interventions2019 · 14,420 citations
- Estimating the reproducibility of psychological scienceAlexander A. Aarts · 2015 · 8,669 citations
- Updated methodological guidance for the conduct of scoping reviewsMicah D.J. Peters · 2020 · 6,688 citations
- Autonomous chemical research with large language modelsDaniil A. Boiko · 2023 · 809 citations
- Role of AI chatbots in education: systematic literature reviewLasha Labadze · 2023 · 791 citations