12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Rethinking Literature Search Evaluation: Deep Research Helps, and Human Citation Lists Are Not a Ground Truth

Gaurav Sahu, Laurent Charlin, Christopher Pal · ArXiv.org · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
2/4
Quality (LMQS)
C
Evidence
0
Citations
0.00
FWCI

Methodology & findings

Study design

Computational pipeline development with empirical evaluation on a benchmark dataset.

Main result

The study demonstrates that a Deep Research pipeline substantially improves literature search recall: "Deep Research raises recall by an order of magnitude over normal search" (Figure 2), increasing recall on ROLLINGEVAL-JUN25 from below 20% to above 80%. Additionally, the authors find critical limitations in using human reference lists as ground truth: "only 51% of human citations are judged moderately relevant or higher, against 86–88% for the strongest AI-based rerankers." The analysis reveals that "humans are 2.5× more likely than the best AI re-rankers to cite a direct collaborator" at distance d=0 (direct co-authorship).

Research paradigm

Empirical/Computational

Author conclusions

The authors conclude that "Our results argue for separating two questions the community usually conflates. Did the system retrieve the papers the authors cited? is a coverage question against a known-imperfect target. Did the system retrieve papers that a neutral reader would judge relevant? is a different, complementary question." They argue that "neither metric is strictly better: recall against human citations captures dimensions an LLM-judge cannot justify (foundational and tool/library attribution, methodological lineage), while the topical rubric captures dimensions that human reference lists undersample (long-tail topically-related work, non-collaborator authors)." Finally, they state that "We argue for reporting both, alongside the OpenAlex-graph diagnostic, rather than treating either in isolation as ground truth."

Risk of bias

Citation bias: Human authors systematically over-cite direct collaborators (2.5× more likely at d=0 co-authorship distance); Domain bias: Benchmark limited to computer science papers from arXiv only; Judge bias: Single LLM-as-a-judge model may have systematic biases in relevance assessment; Selection bias: Bibliography extraction relies on LLM-based PDF parsing which can introduce noise; Data contamination: ROLLINGEVAL-JUN25 chosen to postdate LLM training cutoffs, but guarantee weakens as models update; Sampling bias: Co-authorship graph fetch limited to 1,000 most recent works per author, undercounting distances; Dataset contamination risk mitigated by using June 2025 snapshot post-dating all evaluated LLM training cutoffs; LLM judge bias: single model (GPT-OSS-120B) may have systematic biases in relevance scoring; Co-authorship graph under-counting bias: capping author works at 1,000 most recent papers biases distance estimates conservatively; Selection bias in benchmark: only computer science papers from arXiv; results may not generalize to other domains; Bibliography extraction noise from LLM-based PDF parsing as fallback method; Citation context omission: semantic relevance judge only conditions on title/abstract, missing in-text citation context that may explain relevance; Non-arXiv citations excluded: analysis incomplete for published papers outside arXiv indexing; Single LLM judge for semantic relevance scoring may introduce model-specific bias; Bibliography extraction relies on LLM agents and carries residual noise; Co-authorship graph capping at 1,000 works per author biases distance estimation conservatively; Benchmark limited to arXiv computer science papers; excludes non-arXiv citations; Citation context not modeled; judge conditions only on title and abstract

Limitations

  • The authors explicitly state: "The LLM judge is a single neutral model
  • rubric-conditioned scoring may drift across model and prompt versions, and we do not currently model citation context (in-text spans surrounding the citation)." Additional limitations include: "The co-authorship graph fetch caps each author at 1,000 most recent works, so we under-count distances through hyper-prolific intermediaries, which biases the reported gap in a conservative direction
  • Bibliography extraction relies on LLM agents and can carry residual noise
  • The ROLLINGEVAL benchmark consists only of computer science papers
  • extending it to biomedical, social-science, and humanities papers would test cross-domain generality
  • Finally, the analysis covers only arXiv-indexed references

Open questions raised

  • Need for citation-context modeling: "we do not currently model citation context (in-text spans surrounding the citation)"
  • Cross-domain evaluation: Extending benchmark beyond computer science to biomedical, social-science, and humanities papers
  • Non-arXiv references: Current analysis excludes non-arXiv citations, limiting generalizability
  • Human annotation validation: "Validating against human relevance ratings on a labelled subsample remains a natural next step"
  • Temporal generalization: ROLLINGEVAL-JUN25 contamination-resistance guarantee weakens as models update, requiring periodic refreshes
  • Need for LLM judge validation against human relevance ratings on labelled subsamples
Data: ROLLINGEVAL-JUN25: 250 computer science arXiv papers (to be released); OpenAlex co-authorship graph (publicly available); SemanticScholar API data; arXiv metadata; ROLLINGEVAL-JUN25: 250-paper computer science benchmark (to be released); OpenAlex co-authorship graph (Priem et al., 2022) - public dataset; Semantic Scholar API (Kinney et al., 2023); arXiv indexed papers with metadata; ROLLINGEVAL-JUN25: 250 computer science arXiv papers from June 2025 snapshot (to be released); OpenAlex (CC0 license); SemanticScholar API; arXiv (per-author licenses)Code: Deep Research pipeline code (to be released on Apache 2.0 license); LLM-judge prompts and OpenAlex distance-analysis scripts (to be released); Deep Research pipeline code (to be released under Apache 2.0); LLM-judge prompts (to be released under Apache 2.0); OpenAlex distance-analysis scripts (to be released under Apache 2.0); Deep Research pipeline (to be released under Apache 2.0); LLM-judge prompts (to be released under CC-BY-4.0)Extracted from: pdfAgreement 44%

Explore related topics

Related papers