Towards Self-Evolving Agentic Literature Retrieval
Yuwen Du, Tian Jin, Jing Kang, Xianghe Pang, Jingyi Chai, Tingjia Miao et al. · ArXiv.org · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
Methodology & findings
Study design
Comparative system evaluation using a novel benchmark (PaSaMaster-Bench) with 244 expert-curated tasks spanning 38 scientific disciplines.
Primary method
Design science; agentic AI system design with three architectural layers (Navigator, Librarian Swarm, verified corpus layer)
Main result
PaSaMaster achieves a 16.5× higher F1-score than Google Scholar and a 37.8% higher F1-score than GPT-5.2 at about 1% of the cost, while reducing source hallucination from 32.66% in generative LLMs to zero. Specifically, "PaSaMaster achieves the highest retrieval performance across all main quality metrics, with an NDCG@20 of 39.52, Recall@20 of 33.24, Precision@20 of 23.46, and F1-score@20 of 23.00." The system "achieves 0% hallucination while maintaining the strongest retrieval quality" compared to tool-assisted generative LLMs which "exhibit substantial hallucination rates, including 32.66% for MiniMax-M2.7, 27.54% for Gemini-3.1, 26.59% for Kimi-K2.5."
Research paradigm
Design science; agentic AI systems research
Author conclusions
The authors conclude that "scientific literature discovery can be framed as a Recursive Self-Evolving, evidence-grounded ranking problem rather than as either keyword matching or citation generation." They state that "PaSaMaster shows that Recursive Self-Evolving, evidence-grounded relevance ranking can improve complex intent understanding, eliminate source hallucination and enable low-cost scientific literature discovery at scale." Furthermore, "self-evolving, verified retrieval is a promising foundation for trustworthy AI-assisted science, but its full value will depend on transparent evaluation, continual corpus updating and careful integration into expert research workflows."
Risk of bias
Annotation bias: Expert-curated benchmark may reflect annotator biases and search horizons; Corpus bias: Ranking may inherit biases from the 1.7 billion scientific corpora used; Model bias: Lightweight Ranker model bias; Checklist design bias: Verification checklist construction may not capture all relevant papers; Expert annotation bias in benchmark construction; Disciplinary expertise and search horizon bias of annotators; Potential biases in corpus selection and model training; Checklist design bias in relevance judgments; Annotation bias: Expert curators may have disciplinary biases and limited search horizons; Corpus bias: Biases from scientific literature corpora in relevance scoring; Model bias: Lightweight ranker trained on multidisciplinary data may have domain-specific biases; Checklist design bias: The way evaluation criteria are framed may influence what papers are considered relevant
Limitations
- The authors identify several limitations: "PaSaMaster-Bench is expert-curated, and its target sets may still be influenced by the disciplinary expertise, prior knowledge and search horizon of the annotators." Additionally, "the current evaluation focuses on top-ranked paper retrieval rather than the quality of downstream literature synthesis, hypothesis generation or manuscript writing." Furthermore, "although PaSaMaster eliminates source hallucination by construction, relevance scoring can still inherit biases from corpora, models and checklist design."
Open questions raised
- Evaluation of downstream literature synthesis, hypothesis generation and manuscript writing quality
- Testing whether improved retrieval leads to better scientific outputs in human-in-the-loop settings
- Extension to helping scientists reconstruct historical field development and identify emerging research trajectories
- Generation of candidate research directions grounded in literature
- Transparent evaluation and audit mechanisms for deployment-ready systems
- Continual corpus updating and integration into expert research workflows
Explore related topics
Related papers
- Rayyan—a web and mobile app for systematic reviewsMourad Ouzzani · 2016 · 24,664 citations
- Guidance for conducting systematic scoping reviewsMicah D.J. Peters · 2015 · 7,472 citations
- Which academic search systems are suitable for systematic reviews or meta‐analyses? Evaluating retrieval qualities of Google Scholar, PubMed, and 26 other resourcesMichael Gusenbauer · 2019 · 2,116 citations
- Autonomous chemical research with large language modelsDaniil A. Boiko · 2023 · 809 citations
- Hallucination Rates and Reference Accuracy of ChatGPT and Bard for Systematic Reviews: Comparative AnalysisMikaël Chelli · 2024 · 295 citations
- Artificial intelligence and the conduct of literature reviewsGerit Wagner · 2021 · 275 citations