12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

SciNet: Evaluating AI Agents in Relation-Aware Scientific Literature Retrieval

Chenyang Shao, Yong Li, Fengli Xu · ArXiv.org · 2025

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
3/4
Quality (LMQS)
C
Evidence
0
Citations
0.00
FWCI

Methodology & findings

Study design

Computational benchmark evaluation study.

Main result

Extensive evaluation of three categories of retrieval agents shows that "their accuracy on relation-aware tasks often falls below 20%, highlighting a fundamental shortcoming of current retrieval paradigms." Furthermore, "in a downstream literature review application, agents empowered with SciNet achieve a 25.3% improvement in review quality, highlighting the critical value of relation-aware retrieval for deepening scientific insights." Specific results demonstrate that even top-performing systems achieved only 1.47% recall@50 for novelty and 4.71% for disruption, while path-wise relation identification showed consistency scores of 63.28% but connectivity of only 14.54% for the best model.

Research paradigm

Empirical-computational: systematic evaluation of AI agents through benchmark dataset and comparative performance metrics

Author conclusions

The authors conclude that "capturing literature relations is critical for practical downstream applications. Systems with access to richer relational context produce more coherent, informative, and useful survey reports." They further state that "reconstructing intellectual lineages requires relation-aware retrieval capable of modeling temporal progression and causal reasoning, pushing beyond shallow semantic matching toward genuine knowledge synthesis." Additionally, they emphasize that "providing agents with relational ground truth, such as citation sentiment and network disruption indices (Di), is not merely an enhancement but a necessity for meaningful literature synthesis."

Risk of bias

Selection bias in paper selection: dataset may over-represent highly-cited and well-indexed papers; Recency bias in evaluations: more recent papers may have incomplete citation records; Domain bias: focus on 7 major scientific domains may not represent all scientific fields equally; Language bias: reliance on English-language papers from OpenAlex; Citation bias: using citation networks as ground truth may privilege heavily-cited but not necessarily most novel work; Query construction bias: Standardized test queries derived from structured rules may not reflect real-world user behavior; LLM annotation bias: While achieving 98% agreement with human validation on sentiment classification, LLM-based evaluation could introduce systematic biases; Selection bias in ground truth construction: Manual expert validation at critical nodes could introduce subjective judgment; Temporal bias: Dataset snapshot from July 7, 2025 may not represent complete evolution of scientific fields; Query generation bias: Queries constructed through structured rules and manual review may not reflect authentic researcher information needs; Domain bias: Evaluation limited to 7 scientific domains; generalization to other fields uncertain; Temporal bias: Dataset snapshot from July 2025; future evolution of research patterns not captured; Annotation bias: LLM-based sentiment classification despite high agreement (98%) with human validation, four minor discrepancies remain; Selection bias in ground truth path construction: BFS-based maximum citation count approach may favor highly-cited but not necessarily optimal evolutionary trajectories

Limitations

  • The authors acknowledge that "While this covers 2,640 fine-grained subfields, it does not fully capture the diverse and sometimes unpredictable nature of real-world user queries
  • Future updates will focus on incorporating authentic 'wild queries' from academic search logs and utilizing LLMs for query paraphrasing to better mirror real-world application scenarios." Additionally, the ground truth construction relies on citation counts as a proxy for scientific value, which may not capture all important research trajectories.

Open questions raised

  • Current retrieval agents lack explicit mechanisms to model complex relational networks among scientific papers
  • Existing benchmarks overlook deep relational structures in scientific networks, emphasizing only semantic precision
  • Systems struggle to identify corroborating or conflicting studies
  • Models cannot effectively trace technological lineages and scientific evolution
  • Better mechanisms needed for modeling temporal progression and causal reasoning in citation networks
  • Need for incorporating authentic user queries from academic search logs rather than rule-based synthetic queries
Data: SciNet; OpenAlex; arXiv PDF corpus; SciNet dataset: https://github.com/tsinghua-fib-lab/SciNet (publicly released); OpenAlex meta-database (RELEASE 2025-07-07): https://docs.openalex.org/download-all-data/download-to-your-machine; arXiv PDF corpus (as of July 7, 2025)Code: SciNet Repository; https://github.com/tsinghua-fib-lab/SciNet; SciNetExtracted from: pdfAgreement 48%

Explore related topics

Related papers