SciAtlas: A Large-Scale Knowledge Graph for Automated Scientific Research
Shuofei Qiao, Yunxiang Wei, Jiazheng Fan, Bin Wu, Busheng Zhang, Mengru Wang et al. · ArXiv.org · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
Methodology & findings
Study design
Knowledge graph construction from OpenAlex data sources; neuro-symbolic retrieval algorithm combining lexical matching, vector retrieval, and graph propagation; application demonstrations for downstream tasks including literature review, idea grounding, and research trend prediction..
Main result
SciAtlas integrates "over 43M papers from 26 disciplines, and a total of 157M entities and 3B triplets" to provide "a structured topological cognitive substrate that dismantles disciplinary barriers and furnishes AI agents with a global perspective." The neuro-symbolic retrieval algorithm "achieves a seamless transition from simple semantic matching to deterministic association discovery" and "can be completed within 2 minutes, significantly shorter than LLM-based deep research frameworks, while still delivering high-relevance results with in-depth topological reasoning."
Research paradigm
Pragmatist/Systems Engineering
Author conclusions
"SciAtlas can serve as an effective 'cognitive map' to empower the full loop of automated scientific research while significantly reducing reasoning costs." The authors demonstrate that the system provides "deterministic deep association discovery without requiring frequent iterations of LLMs and high reasoning costs" and present "SciAtlas's capability as a 'cognitive map' to empower the entire loop of automated scientific research" through applications including literature review, idea positioning, research trend synthesis, and academic trajectory exploration.
Risk of bias
Disciplinary bias toward Medicine (18.56%), Social Sciences (10.70%), and Engineering (9.43%), underrepresenting fields like Veterinary (0.16%), Chemistry (1.17%), and others. Language bias: filtering to English-only papers excludes non-English scholarship. Data source bias: dependency on OpenAlex data quality and coverage. Keyword extraction bias: LLM-based extraction may reflect training data biases.; weighting scheme contains multiple manually-tuned hyperparameters that could introduce configuration bias.; Selection bias in paper corpus (English-only papers, papers with abstracts excluded); author name ambiguity not addressed.
Open questions raised
- The authors identify that current academic retrieval tools suffer from:
- organizational deficiency—knowledge scattered in unstructured formats creating "knowledge island" phenomena and impeding interdisciplinary integration
- retrieval paradigm limitations—existing keyword matching and vector-space retrieval lack topological reasoning capabilities
- LLM-based agentic frameworks suffer from high computational costs, latency, and logical hallucinations. Future work mentioned includes progressive opening of hyperparameter configuration for user customization and periodic updates of the knowledge graph.
- organizational form - academic knowledge is scattered in unstructured formats lacking unified paradigms, preventing interdisciplinary integration
Explore related topics
Related papers
- Rayyan—a web and mobile app for systematic reviewsMourad Ouzzani · 2016 · 24,664 citations
- Guidance for conducting systematic scoping reviewsMicah D.J. Peters · 2015 · 7,472 citations
- Which academic search systems are suitable for systematic reviews or meta‐analyses? Evaluating retrieval qualities of Google Scholar, PubMed, and 26 other resourcesMichael Gusenbauer · 2019 · 2,116 citations
- Hallucination Rates and Reference Accuracy of ChatGPT and Bard for Systematic Reviews: Comparative AnalysisMikaël Chelli · 2024 · 295 citations
- Artificial intelligence and the conduct of literature reviewsGerit Wagner · 2021 · 275 citations
- Artificial intelligence for literature reviews: opportunities and challengesF. J. Bolaños · 2024 · 188 citations