12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

SciAtlas: A Large-Scale Knowledge Graph for Automated Scientific Research

Shuofei Qiao, Yunxiang Wei, Jiazheng Fan, Bin Wu, Busheng Zhang, Mengru Wang et al. · ArXiv.org · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
2/4
Quality (LMQS)
C
Evidence
0
Citations
0.00
FWCI

Methodology & findings

Study design

Knowledge graph construction from OpenAlex data sources; neuro-symbolic retrieval algorithm combining lexical matching, vector retrieval, and graph propagation; application demonstrations for downstream tasks including literature review, idea grounding, and research trend prediction..

Main result

SciAtlas integrates "over 43M papers from 26 disciplines, and a total of 157M entities and 3B triplets" to provide "a structured topological cognitive substrate that dismantles disciplinary barriers and furnishes AI agents with a global perspective." The neuro-symbolic retrieval algorithm "achieves a seamless transition from simple semantic matching to deterministic association discovery" and "can be completed within 2 minutes, significantly shorter than LLM-based deep research frameworks, while still delivering high-relevance results with in-depth topological reasoning."

Research paradigm

Pragmatist/Systems Engineering

Author conclusions

"SciAtlas can serve as an effective 'cognitive map' to empower the full loop of automated scientific research while significantly reducing reasoning costs." The authors demonstrate that the system provides "deterministic deep association discovery without requiring frequent iterations of LLMs and high reasoning costs" and present "SciAtlas's capability as a 'cognitive map' to empower the entire loop of automated scientific research" through applications including literature review, idea positioning, research trend synthesis, and academic trajectory exploration.

Risk of bias

Disciplinary bias toward Medicine (18.56%), Social Sciences (10.70%), and Engineering (9.43%), underrepresenting fields like Veterinary (0.16%), Chemistry (1.17%), and others. Language bias: filtering to English-only papers excludes non-English scholarship. Data source bias: dependency on OpenAlex data quality and coverage. Keyword extraction bias: LLM-based extraction may reflect training data biases.; Language bias (non-English papers filtered out); disciplinary representation bias (Medicine 18.56%, Social Sciences 10.70%, with long tail of underrepresented fields); dependency on OpenAlex data quality and coverage; potential bias in LLM-based keyword extraction toward mainstream terminology; weighting scheme contains multiple manually-tuned hyperparameters that could introduce configuration bias.; Selection bias in paper corpus (English-only papers, papers with abstracts excluded); potential coverage bias toward well-documented disciplines in OpenAlex; keyword extraction bias from LLM model choice; author name ambiguity not addressed.

Open questions raised

  • The authors identify that current academic retrieval tools suffer from: (1) organizational deficiency—knowledge scattered in unstructured formats creating "knowledge island" phenomena and impeding interdisciplinary integration; (2) retrieval paradigm limitations—existing keyword matching and vector-space retrieval lack topological reasoning capabilities; (3) LLM-based agentic frameworks suffer from high computational costs, latency, and logical hallucinations. Future work mentioned includes progressive opening of hyperparameter configuration for user customization and periodic updates of the knowledge graph.
  • The authors identify that current academic retrieval tools suffer from two major issues: (1) organizational form - academic knowledge is scattered in unstructured formats lacking unified paradigms, preventing interdisciplinary integration; (2) retrieval paradigm - existing tools rely on superficial keyword matching or vector-space retrieval without genuine topological reasoning, and deep-research frameworks incur high computational costs and are susceptible to logical hallucinations. They position SciAtlas as addressing the need for structured knowledge organization and deterministic reasoning to support the automated scientific research loop.
  • The authors identify the need for better academic knowledge organization to overcome the "knowledge island" phenomenon and the limitations of existing retrieval tools that rely on superficial keyword matching or vector-space semantic retrieval lacking "topological reasoning capabilities." They note that agentic deep-research frameworks are "prone to logical hallucinations and consuming high inference costs."
Data: SciAtlas is built on OpenAlex (https://openalex.org/), containing over 480 million academic publications. The authors state: "We have released the interfaces for KG retrieval and various downstream tasks in our GitHub repo" (https://github.com/zjunlp/SciAtlas).; SciAtlas knowledge graph interfaces and downstream task implementations available at: https://github.com/zjunlp/SciAtlas; Source data from OpenAlex (https://openalex.org/), a fully open-source library of scholarly resources encompassing over 480 million academic publications.; SciAtlas knowledge graph (interfaces available via GitHub: https://github.com/zjunlp/SciAtlas); OpenAlex library (https://openalex.org/)Code: https://github.com/zjunlp/SciAtlas - GitHub repository for SciAtlas interfaces and downstream tasks. Code for KG construction is promised to be open-sourced to support evolution.; https://github.com/zjunlp/SciAtlas (interfaces for KG retrieval and various downstream tasks); GROBID for PDF metadata extraction (https://github.com/grobidOrg/grobid); bge-large-en-v1.5 embedding model referenced; Neo4j database system for deployment.; https://github.com/zjunlp/SciAtlas; GROBID (https://github.com/grobidOrg/grobid); SciGraph project (http://scigraph.openkg.cn/)Extracted from: pdfAgreement 50%

Explore related topics

Related papers