12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Deep Research: A Systematic Survey

Yiqun Chen, Haitao Li, Weiwei Sun, Shaoqing Ni, Yougang Lyu, Run-Ze Fan et al. · Preprints.org · 2025

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
2/4
Quality (LMQS)
I
Evidence
0
Citations

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.20944/preprints202511.2077.v1

Methodology & findings

Study design

Systematic narrative literature review - The authors conducted a comprehensive synthesis of existing Deep Research literature, mapping core components to representative system implementations, consolidating key techniques and evaluation methodologies, and establishing a foundation for consistent benchmarking.

Main result

The survey identifies that "Deep Research (DR) frames LLMs within an end-to-end research workflow that iteratively decomposes complex problems, acquire evidence via tool use, and synthesizes validated insights into coherent long-form answers." The paper proposes "a three-stage roadmap for DR systems, illustrating their broad applications ranging from agentic information seeking to autonomous scientific discovery" and establishes "four foundational components in DR: (i) query planning, which decomposes the initially input query into a series of simpler, sub-queries; (ii) information acquisition, which invokes external retrieval, web browsing, or various tools on demand; (iii) memory management, which ensures relevant task-solving context through controlled updating or folding; (iv) answer generation, which produces comprehensive outputs with explicit source attribution."

Research paradigm

Interpretive/Descriptive - Systematic survey of literature on Deep Research systems without quantitative synthesis or meta-analysis

Author conclusions

The authors conclude that "Deep research (DR) stands at the frontier of transforming large language models from passive responders into autonomous investigators capable of iterative reasoning, evidence synthesis, and verifiable knowledge creation. This survey consolidates recent advances in architectures, optimization methods, and evaluation frameworks, providing a unified roadmap for understanding and building future DR systems." They further state: "By investigating relevant works, this survey facilitates future research and accelerates the advancement of DR systems toward more general, reliable, and interpretable intelligence. Given the rapid evolution of this field, we will continuously update this survey to encompass emerging paradigms such as multimodal reasoning, self-evolving memory, and agentic reinforcement learning."

Risk of bias

Selection bias in literature coverage - survey may emphasize publicly available or English-language publications; Publication bias - tendency to include successful or well-documented system implementations over failed or negative cases; Recency bias - rapid evolution of the field may lead to outdated information even during review process; Author affiliation bias - representation of systems from major research institutions may be overrepresented

Limitations

  • The paper acknowledges that "Despite rapid progress, there remains no comprehensive survey that systematically analyzes the key components, technical details, and open challenges of DR." Furthermore, the authors note that "reliably evaluating model-generated long-form outputs, especially research-style reports in response to open-ended and high-level queries, remains an open and pressing challenge" and that "most existing approaches rely on LLM-as-a-Judge to directly evaluate general dimensions such as content factuality, structural coherence, and readability while ignoring crucial dimensions." Additionally, the survey identifies that "existing DR systems, such as Search-R1, rely too heavily on answer correctness to guide the entire search pipeline and lack fine-grained guidance on when to retrieve, leading to both over-retrieval and under-retrieval."

Open questions raised

  • Lack of fine-grained guidance on when to retrieve in existing DR systems, leading to both over-retrieval and under-retrieval
  • Challenges in reliably evaluating model-generated long-form outputs, especially research-style reports in response to open-ended queries
  • Logical evaluation across multiple granularities for long-form coherence
  • Distinction between novelty and hallucination in DR outputs
  • Bias and efficiency issues in LLM-as-Judge evaluation approaches
  • Training stability and instability in multi-turn RL for DR systems
Data: NQ (Natural Questions); TriviaQA; SimpleQA; HotpotQA; 2WikiMultihopQA; Bamboogle; MuSiQue; FRAMES; GPQA; GAIA; HLE; InfoDeepSeek; AssistantBench; Mind2Web; Mind2Web 2; BrowseComp; DeepResearchGym; WebArena; WebWalkerQA; WideSearch; MMInA; AutoSurvey; ReportBench; SurveyGen; Deep Research Comparator; DeepResearch Bench; ResearcherBench; LiveDRBench; PROXYQA; SCHOLARQABENCH; Paper2Poster; PosterGen; P2PInstruct; Doc2PPT; SLIDESBENCH; Zenodo10K; TSBench; AI Idea Bench 2025; Scientist-Bench; PaperBench; ASAP-Review; REVIEW-5k; REVIEWER2; DeepReview-Bench; SWE-Bench; ScienceWorld; GPT-Simulator; DiscoveryWorld; CORE-Bench; MLE; RE-Bench; DSBench; Spider2-V; DSEval; UnivEARTH; Commit0; NQ (Natural Questions) - https://github.com/google-research-datasets/natural-questions; MultiHop-RAG; GPQA (Graduate-level Google-Proof QA); HLE (Highly complex Long-form Evaluation); DeepResearch Comparator; MLE (Machine Learning Engineering)Extracted from: pdfAgreement 83%

Explore related topics

Related papers