12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Cited but Not Verified: Parsing and Evaluating Source Attribution in LLM Deep Research Agents

Hailey Onweller, Elias Lumer, Austin Huber, Pia Ramchandani, Vamse Kumar Subbiah, Corey Feld · ArXiv.org · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
1/4
Quality (LMQS)
C
Evidence
0
Citations
0.00
FWCI

Methodology & findings

Study design

Computational framework with three-stage evaluation pipeline: (1) Markdown AST parser for citation-claim extraction, (2) web content retrieval, (3) multi-dimensional evaluation across Link Works (URL accessibility), Relevant Content (topical relevance via LLM-as-a-judge), and Fact Check (factual accuracy via LLM-as-a-judge).

Main result

The study reveals that "even the strongest frontier models maintain link validity above 94% and content relevance above 80%, yet achieve only 39-77% factual accuracy, exposing a critical disconnect between surface-level citation quality and factual reliability." Additionally, "Fact Check accuracy drops approximately 42% on average as search depth scales from 2 to 150 tool calls, while Link Works and Relevant Content remain stable, suggesting that information overload impairs factual synthesis rather than improving it."

Research paradigm

empirical/computational evaluation

Author conclusions

"Our evaluation of 14 LLMs reveals a critical disconnect between surface-level citation quality and factual reliability. Models consistently produce working links to relevant pages, yet factual accuracy remains the weakest dimension across all providers." The authors further conclude: "We further demonstrate an information overload effect, where increased search depth degrades factual accuracy while surface metrics remain stable, providing evidence that more retrieval does not produce more accurate citations."

Risk of bias

LLM-as-a-judge biases (position bias, verbosity bias, self-enhancement effects); Temporal instability of web citations affecting reproducibility; Selection bias in query set (130 queries from two specific benchmarks); Calibration bias: Fact Check evaluator calibrated through only 50-100 human reviews; Model-specific biases in evaluation judge; LLM judge bias (position bias, self-enhancement effects) despite human calibration; Limited scope to web-search-capable models, excluding enterprise RAG systems; Human review calibration limited to 50-100 judgments for Fact Check evaluator; Potential biases in selection of research queries from DeepResearch Bench and BrowseComp; LLM judge bias (position bias, self-enhancement effects); Potential selection bias in research query dataset (drawn from DeepResearch Bench and BrowseComp); Human calibration bias in Fact Check evaluator training (50-100 judgments only)

Limitations

  • The authors state: "the LLM-as-a-judge approach for Relevant Content and Fact Check evaluations, despite calibration through human review, may retain biases inherent to the judge model, including position bias and self-enhancement effects." Additionally, "web citations are temporally unstable
  • URLs that were accessible during evaluation may become unavailable due to content removal, domain expiration, or access policy changes, and source content itself may change." Finally, "the evaluation is limited to models with web search capabilities, excluding enterprise RAG deployments that cite internal document corpora."

Open questions raised

  • No end-to-end framework combining citation extraction with multi-dimensional quality assessment across URL accessibility, topical relevance, and factual accuracy
  • No systematic comparison across major LLM providers in deep research settings
  • Relationship between search depth and citation quality remains unexplored
  • Evaluation limited to web search capabilities; enterprise RAG deployments with internal document corpora not covered
  • Need for longitudinal studies tracking citation persistence over time
  • Need for domain-specific research task evaluation
Data: DeepResearch Bench (Du et al., 2025); BrowseComp (Wei et al., 2025); 130 research queries drawn from DeepResearch Bench (Du et al., 2025); Research queries from BrowseComp (Wei et al., 2025)Extracted from: pdfAgreement 66%

Explore related topics

Related papers