SciAtlas: A Large-Scale Knowledge Graph for Automated Scientific Research
Shuofei Qiao, Yunxiang Wei, Jiazheng Fan, Bin Wu, Busheng Zhang, Mengru Wang et al. · ArXiv.org · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
Methodology & findings
Study design
Knowledge graph construction from OpenAlex data sources; neuro-symbolic retrieval algorithm combining lexical matching, vector retrieval, and graph propagation; application demonstrations for downstream tasks including literature review, idea grounding, and research trend prediction..
Main result
SciAtlas integrates "over 43M papers from 26 disciplines, and a total of 157M entities and 3B triplets" to provide "a structured topological cognitive substrate that dismantles disciplinary barriers and furnishes AI agents with a global perspective." The neuro-symbolic retrieval algorithm "achieves a seamless transition from simple semantic matching to deterministic association discovery" and "can be completed within 2 minutes, significantly shorter than LLM-based deep research frameworks, while still delivering high-relevance results with in-depth topological reasoning."
Research paradigm
Pragmatist/Systems Engineering
Author conclusions
"SciAtlas can serve as an effective 'cognitive map' to empower the full loop of automated scientific research while significantly reducing reasoning costs." The authors demonstrate that the system provides "deterministic deep association discovery without requiring frequent iterations of LLMs and high reasoning costs" and present "SciAtlas's capability as a 'cognitive map' to empower the entire loop of automated scientific research" through applications including literature review, idea positioning, research trend synthesis, and academic trajectory exploration.
Risk of bias
Disciplinary bias toward Medicine (18.56%), Social Sciences (10.70%), and Engineering (9.43%), underrepresenting fields like Veterinary (0.16%), Chemistry (1.17%), and others. Language bias: filtering to English-only papers excludes non-English scholarship. Data source bias: dependency on OpenAlex data quality and coverage. Keyword extraction bias: LLM-based extraction may reflect training data biases.; Language bias (non-English papers filtered out); disciplinary representation bias (Medicine 18.56%, Social Sciences 10.70%, with long tail of underrepresented fields); dependency on OpenAlex data quality and coverage; potential bias in LLM-based keyword extraction toward mainstream terminology; weighting scheme contains multiple manually-tuned hyperparameters that could introduce configuration bias.; Selection bias in paper corpus (English-only papers, papers with abstracts excluded); potential coverage bias toward well-documented disciplines in OpenAlex; keyword extraction bias from LLM model choice; author name ambiguity not addressed.
Open questions raised
- The authors identify that current academic retrieval tools suffer from: (1) organizational deficiency—knowledge scattered in unstructured formats creating "knowledge island" phenomena and impeding interdisciplinary integration; (2) retrieval paradigm limitations—existing keyword matching and vector-space retrieval lack topological reasoning capabilities; (3) LLM-based agentic frameworks suffer from high computational costs, latency, and logical hallucinations. Future work mentioned includes progressive opening of hyperparameter configuration for user customization and periodic updates of the knowledge graph.
- The authors identify that current academic retrieval tools suffer from two major issues: (1) organizational form - academic knowledge is scattered in unstructured formats lacking unified paradigms, preventing interdisciplinary integration; (2) retrieval paradigm - existing tools rely on superficial keyword matching or vector-space retrieval without genuine topological reasoning, and deep-research frameworks incur high computational costs and are susceptible to logical hallucinations. They position SciAtlas as addressing the need for structured knowledge organization and deterministic reasoning to support the automated scientific research loop.
- The authors identify the need for better academic knowledge organization to overcome the "knowledge island" phenomenon and the limitations of existing retrieval tools that rely on superficial keyword matching or vector-space semantic retrieval lacking "topological reasoning capabilities." They note that agentic deep-research frameworks are "prone to logical hallucinations and consuming high inference costs."
Explore related topics
Related papers
- Rayyan—a web and mobile app for systematic reviewsMourad Ouzzani · 2016 · 24,664 citations
- Guidance for conducting systematic scoping reviewsMicah D.J. Peters · 2015 · 7,472 citations
- Which academic search systems are suitable for systematic reviews or meta‐analyses? Evaluating retrieval qualities of Google Scholar, PubMed, and 26 other resourcesMichael Gusenbauer · 2019 · 2,116 citations
- Hallucination Rates and Reference Accuracy of ChatGPT and Bard for Systematic Reviews: Comparative AnalysisMikaël Chelli · 2024 · 295 citations
- Artificial intelligence and the conduct of literature reviewsGerit Wagner · 2021 · 275 citations
- Artificial intelligence for literature reviews: opportunities and challengesF. J. Bolaños · 2024 · 188 citations