SciNet: Evaluating AI Agents in Relation-Aware Scientific Literature Retrieval
Chenyang Shao, Yong Li, Fengli Xu · ArXiv.org · 2025
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
Methodology & findings
Study design
Computational benchmark evaluation study.
Main result
Extensive evaluation of three categories of retrieval agents shows that "their accuracy on relation-aware tasks often falls below 20%, highlighting a fundamental shortcoming of current retrieval paradigms." Furthermore, "in a downstream literature review application, agents empowered with SciNet achieve a 25.3% improvement in review quality, highlighting the critical value of relation-aware retrieval for deepening scientific insights." Specific results demonstrate that even top-performing systems achieved only 1.47% recall@50 for novelty and 4.71% for disruption, while path-wise relation identification showed consistency scores of 63.28% but connectivity of only 14.54% for the best model.
Research paradigm
Empirical-computational: systematic evaluation of AI agents through benchmark dataset and comparative performance metrics
Author conclusions
The authors conclude that "capturing literature relations is critical for practical downstream applications. Systems with access to richer relational context produce more coherent, informative, and useful survey reports." They further state that "reconstructing intellectual lineages requires relation-aware retrieval capable of modeling temporal progression and causal reasoning, pushing beyond shallow semantic matching toward genuine knowledge synthesis." Additionally, they emphasize that "providing agents with relational ground truth, such as citation sentiment and network disruption indices (Di), is not merely an enhancement but a necessity for meaningful literature synthesis."
Risk of bias
Selection bias in paper selection: dataset may over-represent highly-cited and well-indexed papers; Recency bias in evaluations: more recent papers may have incomplete citation records; Domain bias: focus on 7 major scientific domains may not represent all scientific fields equally; Language bias: reliance on English-language papers from OpenAlex; Citation bias: using citation networks as ground truth may privilege heavily-cited but not necessarily most novel work; Query construction bias: Standardized test queries derived from structured rules may not reflect real-world user behavior; LLM annotation bias: While achieving 98% agreement with human validation on sentiment classification, LLM-based evaluation could introduce systematic biases; Selection bias in ground truth construction: Manual expert validation at critical nodes could introduce subjective judgment; Temporal bias: Dataset snapshot from July 7, 2025 may not represent complete evolution of scientific fields; Query generation bias: Queries constructed through structured rules and manual review may not reflect authentic researcher information needs; Domain bias: Evaluation limited to 7 scientific domains; generalization to other fields uncertain; Temporal bias: Dataset snapshot from July 2025; future evolution of research patterns not captured; Annotation bias: LLM-based sentiment classification despite high agreement (98%) with human validation, four minor discrepancies remain; Selection bias in ground truth path construction: BFS-based maximum citation count approach may favor highly-cited but not necessarily optimal evolutionary trajectories
Limitations
- The authors acknowledge that "While this covers 2,640 fine-grained subfields, it does not fully capture the diverse and sometimes unpredictable nature of real-world user queries
- Future updates will focus on incorporating authentic 'wild queries' from academic search logs and utilizing LLMs for query paraphrasing to better mirror real-world application scenarios." Additionally, the ground truth construction relies on citation counts as a proxy for scientific value, which may not capture all important research trajectories.
Open questions raised
- Current retrieval agents lack explicit mechanisms to model complex relational networks among scientific papers
- Existing benchmarks overlook deep relational structures in scientific networks, emphasizing only semantic precision
- Systems struggle to identify corroborating or conflicting studies
- Models cannot effectively trace technological lineages and scientific evolution
- Better mechanisms needed for modeling temporal progression and causal reasoning in citation networks
- Need for incorporating authentic user queries from academic search logs rather than rule-based synthetic queries
Explore related topics
Related papers
- Rayyan—a web and mobile app for systematic reviewsMourad Ouzzani · 2016 · 24,664 citations
- Guidance for conducting systematic scoping reviewsMicah D.J. Peters · 2015 · 7,472 citations
- Which academic search systems are suitable for systematic reviews or meta‐analyses? Evaluating retrieval qualities of Google Scholar, PubMed, and 26 other resourcesMichael Gusenbauer · 2019 · 2,116 citations
- Autonomous chemical research with large language modelsDaniil A. Boiko · 2023 · 809 citations
- Hallucination Rates and Reference Accuracy of ChatGPT and Bard for Systematic Reviews: Comparative AnalysisMikaël Chelli · 2024 · 295 citations
- Artificial intelligence and the conduct of literature reviewsGerit Wagner · 2021 · 275 citations