Cited but Not Verified: Parsing and Evaluating Source Attribution in LLM Deep Research Agents
Hailey Onweller, Elias Lumer, Austin Huber, Pia Ramchandani, Vamse Kumar Subbiah, Corey Feld · ArXiv.org · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
Methodology & findings
Study design
Computational framework with three-stage evaluation pipeline: (1) Markdown AST parser for citation-claim extraction, (2) web content retrieval, (3) multi-dimensional evaluation across Link Works (URL accessibility), Relevant Content (topical relevance via LLM-as-a-judge), and Fact Check (factual accuracy via LLM-as-a-judge).
Main result
The study reveals that "even the strongest frontier models maintain link validity above 94% and content relevance above 80%, yet achieve only 39-77% factual accuracy, exposing a critical disconnect between surface-level citation quality and factual reliability." Additionally, "Fact Check accuracy drops approximately 42% on average as search depth scales from 2 to 150 tool calls, while Link Works and Relevant Content remain stable, suggesting that information overload impairs factual synthesis rather than improving it."
Research paradigm
empirical/computational evaluation
Author conclusions
"Our evaluation of 14 LLMs reveals a critical disconnect between surface-level citation quality and factual reliability. Models consistently produce working links to relevant pages, yet factual accuracy remains the weakest dimension across all providers." The authors further conclude: "We further demonstrate an information overload effect, where increased search depth degrades factual accuracy while surface metrics remain stable, providing evidence that more retrieval does not produce more accurate citations."
Risk of bias
LLM-as-a-judge biases (position bias, verbosity bias, self-enhancement effects); Temporal instability of web citations affecting reproducibility; Selection bias in query set (130 queries from two specific benchmarks); Calibration bias: Fact Check evaluator calibrated through only 50-100 human reviews; Model-specific biases in evaluation judge; LLM judge bias (position bias, self-enhancement effects) despite human calibration; Limited scope to web-search-capable models, excluding enterprise RAG systems; Human review calibration limited to 50-100 judgments for Fact Check evaluator; Potential biases in selection of research queries from DeepResearch Bench and BrowseComp; LLM judge bias (position bias, self-enhancement effects); Potential selection bias in research query dataset (drawn from DeepResearch Bench and BrowseComp); Human calibration bias in Fact Check evaluator training (50-100 judgments only)
Limitations
- The authors state: "the LLM-as-a-judge approach for Relevant Content and Fact Check evaluations, despite calibration through human review, may retain biases inherent to the judge model, including position bias and self-enhancement effects." Additionally, "web citations are temporally unstable
- URLs that were accessible during evaluation may become unavailable due to content removal, domain expiration, or access policy changes, and source content itself may change." Finally, "the evaluation is limited to models with web search capabilities, excluding enterprise RAG deployments that cite internal document corpora."
Open questions raised
- No end-to-end framework combining citation extraction with multi-dimensional quality assessment across URL accessibility, topical relevance, and factual accuracy
- No systematic comparison across major LLM providers in deep research settings
- Relationship between search depth and citation quality remains unexplored
- Evaluation limited to web search capabilities; enterprise RAG deployments with internal document corpora not covered
- Need for longitudinal studies tracking citation persistence over time
- Need for domain-specific research task evaluation
Explore related topics
Related papers
- Rayyan—a web and mobile app for systematic reviewsMourad Ouzzani · 2016 · 24,664 citations
- Guidance for conducting systematic scoping reviewsMicah D.J. Peters · 2015 · 7,472 citations
- Which academic search systems are suitable for systematic reviews or meta‐analyses? Evaluating retrieval qualities of Google Scholar, PubMed, and 26 other resourcesMichael Gusenbauer · 2019 · 2,116 citations
- Comparing scientific abstracts generated by ChatGPT to real abstracts with detectors and blinded human reviewersCatherine A. Gao · 2023 · 657 citations
- Fabrication and errors in the bibliographic citations generated by ChatGPTWilliam H. Walters · 2023 · 352 citations
- Hallucination Rates and Reference Accuracy of ChatGPT and Bard for Systematic Reviews: Comparative AnalysisMikaël Chelli · 2024 · 295 citations