Rethinking Literature Search Evaluation: Deep Research Helps, and Human Citation Lists Are Not a Ground Truth
Gaurav Sahu, Laurent Charlin, Christopher Pal · ArXiv.org · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
Methodology & findings
Study design
Computational pipeline development with empirical evaluation on a benchmark dataset.
Main result
The study demonstrates that a Deep Research pipeline substantially improves literature search recall: "Deep Research raises recall by an order of magnitude over normal search" (Figure 2), increasing recall on ROLLINGEVAL-JUN25 from below 20% to above 80%. Additionally, the authors find critical limitations in using human reference lists as ground truth: "only 51% of human citations are judged moderately relevant or higher, against 86–88% for the strongest AI-based rerankers." The analysis reveals that "humans are 2.5× more likely than the best AI re-rankers to cite a direct collaborator" at distance d=0 (direct co-authorship).
Research paradigm
Empirical/Computational
Author conclusions
The authors conclude that "Our results argue for separating two questions the community usually conflates. Did the system retrieve the papers the authors cited? is a coverage question against a known-imperfect target. Did the system retrieve papers that a neutral reader would judge relevant? is a different, complementary question." They argue that "neither metric is strictly better: recall against human citations captures dimensions an LLM-judge cannot justify (foundational and tool/library attribution, methodological lineage), while the topical rubric captures dimensions that human reference lists undersample (long-tail topically-related work, non-collaborator authors)." Finally, they state that "We argue for reporting both, alongside the OpenAlex-graph diagnostic, rather than treating either in isolation as ground truth."
Risk of bias
Citation bias: Human authors systematically over-cite direct collaborators (2.5× more likely at d=0 co-authorship distance); Domain bias: Benchmark limited to computer science papers from arXiv only; Judge bias: Single LLM-as-a-judge model may have systematic biases in relevance assessment; Selection bias: Bibliography extraction relies on LLM-based PDF parsing which can introduce noise; Data contamination: ROLLINGEVAL-JUN25 chosen to postdate LLM training cutoffs, but guarantee weakens as models update; Sampling bias: Co-authorship graph fetch limited to 1,000 most recent works per author, undercounting distances; Dataset contamination risk mitigated by using June 2025 snapshot post-dating all evaluated LLM training cutoffs; LLM judge bias: single model (GPT-OSS-120B) may have systematic biases in relevance scoring; Co-authorship graph under-counting bias: capping author works at 1,000 most recent papers biases distance estimates conservatively; Selection bias in benchmark: only computer science papers from arXiv; results may not generalize to other domains; Bibliography extraction noise from LLM-based PDF parsing as fallback method; Citation context omission: semantic relevance judge only conditions on title/abstract, missing in-text citation context that may explain relevance; Non-arXiv citations excluded: analysis incomplete for published papers outside arXiv indexing; Single LLM judge for semantic relevance scoring may introduce model-specific bias; Bibliography extraction relies on LLM agents and carries residual noise; Co-authorship graph capping at 1,000 works per author biases distance estimation conservatively; Benchmark limited to arXiv computer science papers; excludes non-arXiv citations; Citation context not modeled; judge conditions only on title and abstract
Limitations
- The authors explicitly state: "The LLM judge is a single neutral model
- rubric-conditioned scoring may drift across model and prompt versions, and we do not currently model citation context (in-text spans surrounding the citation)." Additional limitations include: "The co-authorship graph fetch caps each author at 1,000 most recent works, so we under-count distances through hyper-prolific intermediaries, which biases the reported gap in a conservative direction
- Bibliography extraction relies on LLM agents and can carry residual noise
- The ROLLINGEVAL benchmark consists only of computer science papers
- extending it to biomedical, social-science, and humanities papers would test cross-domain generality
- Finally, the analysis covers only arXiv-indexed references
Open questions raised
- Need for citation-context modeling: "we do not currently model citation context (in-text spans surrounding the citation)"
- Cross-domain evaluation: Extending benchmark beyond computer science to biomedical, social-science, and humanities papers
- Non-arXiv references: Current analysis excludes non-arXiv citations, limiting generalizability
- Human annotation validation: "Validating against human relevance ratings on a labelled subsample remains a natural next step"
- Temporal generalization: ROLLINGEVAL-JUN25 contamination-resistance guarantee weakens as models update, requiring periodic refreshes
- Need for LLM judge validation against human relevance ratings on labelled subsamples
Explore related topics
Related papers
- Rayyan—a web and mobile app for systematic reviewsMourad Ouzzani · 2016 · 24,664 citations
- Guidance for conducting systematic scoping reviewsMicah D.J. Peters · 2015 · 7,472 citations
- Which academic search systems are suitable for systematic reviews or meta‐analyses? Evaluating retrieval qualities of Google Scholar, PubMed, and 26 other resourcesMichael Gusenbauer · 2019 · 2,116 citations
- Autonomous chemical research with large language modelsDaniil A. Boiko · 2023 · 809 citations
- Comparing scientific abstracts generated by ChatGPT to real abstracts with detectors and blinded human reviewersCatherine A. Gao · 2023 · 657 citations
- Fabrication and errors in the bibliographic citations generated by ChatGPTWilliam H. Walters · 2023 · 352 citations