Harnessing AtomisticSkills for Agentic Atomistic Research
Bowen Deng, Bohan Li, Matthew Cox, Hoje Chun, Juno Nam, Artur Lyssenko et al. · arXiv (Cornell University) · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
Methodology & findings
Study design
Design science approach combining: (1) systematic coverage analysis via text-mining 500 published papers from npj Computational Materials, The Journal of Physical Chemistry Letters, and PubMed/arXiv for drug discovery; (2) hierarchical decomposition of scientific workflows into skills and tools; (3) integration of >100 human-curated multidisciplinary skills with MCP (Model Context Protocol) toolboxes; (4) case study demonstrations across six scientific domains; (5) validation against literature-reported values..
Primary method
Design science; hierarchical decomposition of scientific knowledge into workflows, skills, and tools; agile harness engineering approach building on top of general-purpose coding agents
Main result
AtomisticSkills demonstrates comprehensive agentic research infrastructure with "more than 100 human-curated multidisciplinary skills, including database access, thermodynamics and kinetics modeling, and diverse simulation engines employing machine learning interatomic potentials (MLIPs) and density functional theory (DFT)." The framework achieves "an average of 56.2% coverage" of computational materials skills used in literature, with "approximately 15% of the evaluated literature reaching 100% operational mapping." The paper validates functional coverage through multiple case studies including generative design of Li-ion solid-state electrolytes, high-throughput MOF screening, MLIP benchmarking, and structure-based virtual screening for drug discovery.
Research paradigm
Design science / Engineering
Author conclusions
The authors conclude that "AtomisticSkills provides a critical agent infrastructure towards building fully autonomous AI scientists." They emphasize that "by hierarchically decomposing scientific workflows into agent skills and tools, AtomisticSkills provides agents with modular, extensible, and plug-and-play research capabilities." Furthermore, they note that "a highly beneficial way to shape the outer harness is via the construction of a library of domain-specific agent skills," and that "this hierarchical decomposition empowers the agent to build complex scientific workflows dynamically from tested skills and tools, without sacrificing the deterministic reliability required for rigorous atomistic research."
Risk of bias
Selection bias in literature mining: 500 papers sampled from specific journals (npj Computational Materials, Journal of Physical Chemistry Letters, PubMed/arXiv) may not represent full landscape of atomistic research; Skill identification bias: LLM-assisted skill extraction and categorization may introduce inconsistencies; Coverage analysis methodology: External LLM (GPT-5.4-mini) used for skill identification, results reviewed by human experts—potential for subjective classification; Case study selection: Six case studies chosen to demonstrate capabilities; unclear if representative of real-world usage patterns; Selection bias in literature sampling (500 papers from specific journals may not represent all computational materials research); Coverage analysis depends on LLM skill identification accuracy (GPT-5.4-mini used for decomposition); Funding bias: Shell Inc. and other corporate sponsors funded portions of the research; Selection bias in literature mining: 500 papers sampled from specific journals may not represent all computational materials science research; Coverage analysis dependent on LLM-based skill identification (GPT-5.4-mini), which may introduce hallucination or misclassification; Funding bias: Research supported by Shell Inc., which may influence scope or validation choices; Validation limited to post-hoc literature comparison rather than prospective user studies
Limitations
- The framework has notable limitations: "Notably, there are around 12% of journal articles with 0 coverage
- These missing coverages mainly correspond to articles with novel algorithmic developments, experimental methods, and special domains that require less frequent computational skills." Additionally, the authors note that "the simulated conductivities serve strictly as literature-standard computational screening metrics rather than rigorously converged values targeted for direct experimental validation."
Open questions raised
- The paper identifies that "an open, systematic, and extensible skill framework dedicated to atomistic materials science and computational chemistry remains a critical open challenge" prior to this work. Future directions implicitly include extending coverage to novel algorithmic developments, experimental methods, and special domains with less frequent computational skills (currently 12% coverage gap). The work also suggests extending framework capabilities to additional scientific domains and improving harmonization with evolving general-purpose coding agents.
- The authors identify that "an open, systematic, and extensible skill framework dedicated to atomistic materials science and computational chemistry remains a critical open challenge." They note that existing agentic infrastructure has bottlenecks: "Most scientific agents were built as monolithic, standalone platforms atop orchestration frameworks such as LangChain, consequently, developers of these platforms are forced to reinvent fundamental software infrastructure." The gap addressed is bridging deterministic reliability of traditional automation with flexible agentic reasoning.
- The paper identifies that "an open, systematic, and extensible skill framework dedicated to atomistic materials science and computational chemistry remains a critical open challenge." Additional gaps include: (1) lack of harness engineering approaches layered on general-purpose agents rather than monolithic platforms; (2) limited skill libraries for computational sciences beyond bioinformatics; (3) difficulty scaling LLM agents to manage rigor and complexity of atomistic research.
Explore related topics
Related papers
- The effects of over-reliance on AI dialogue systems on students' cognitive abilities: a systematic reviewChunpeng Zhai · 2024 · 1,009 citations
- Autonomous chemical research with large language modelsDaniil A. Boiko · 2023 · 809 citations
- Teacher support and student motivation to learn with Artificial Intelligence (AI) based chatbotThomas K. F. Chiu · 2023 · 617 citations
- Beware of metacognitive laziness: Effects of generative artificial intelligence on learning motivation, processes, and performanceYizhou Fan · 2024 · 419 citations
- What ChatGPT means for universities: Perceptions of scholars and studentsMehmet Fırat · 2023 · 405 citations
- Impact of AI assistance on student agencyAli Darvishi · 2023 · 384 citations