12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Harnessing AtomisticSkills for Agentic Atomistic Research

Bowen Deng, Bohan Li, Matthew Cox, Hoje Chun, Juno Nam, Artur Lyssenko et al. · arXiv (Cornell University) · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
2/4
Quality (LMQS)
D
Evidence
0
Citations
0.00
FWCI

Methodology & findings

Study design

Design science approach combining: (1) systematic coverage analysis via text-mining 500 published papers from npj Computational Materials, The Journal of Physical Chemistry Letters, and PubMed/arXiv for drug discovery; (2) hierarchical decomposition of scientific workflows into skills and tools; (3) integration of >100 human-curated multidisciplinary skills with MCP (Model Context Protocol) toolboxes; (4) case study demonstrations across six scientific domains; (5) validation against literature-reported values..

Primary method

Design science; hierarchical decomposition of scientific knowledge into workflows, skills, and tools; agile harness engineering approach building on top of general-purpose coding agents

Main result

AtomisticSkills demonstrates comprehensive agentic research infrastructure with "more than 100 human-curated multidisciplinary skills, including database access, thermodynamics and kinetics modeling, and diverse simulation engines employing machine learning interatomic potentials (MLIPs) and density functional theory (DFT)." The framework achieves "an average of 56.2% coverage" of computational materials skills used in literature, with "approximately 15% of the evaluated literature reaching 100% operational mapping." The paper validates functional coverage through multiple case studies including generative design of Li-ion solid-state electrolytes, high-throughput MOF screening, MLIP benchmarking, and structure-based virtual screening for drug discovery.

Research paradigm

Design science / Engineering

Author conclusions

The authors conclude that "AtomisticSkills provides a critical agent infrastructure towards building fully autonomous AI scientists." They emphasize that "by hierarchically decomposing scientific workflows into agent skills and tools, AtomisticSkills provides agents with modular, extensible, and plug-and-play research capabilities." Furthermore, they note that "a highly beneficial way to shape the outer harness is via the construction of a library of domain-specific agent skills," and that "this hierarchical decomposition empowers the agent to build complex scientific workflows dynamically from tested skills and tools, without sacrificing the deterministic reliability required for rigorous atomistic research."

Risk of bias

Selection bias in literature mining: 500 papers sampled from specific journals (npj Computational Materials, Journal of Physical Chemistry Letters, PubMed/arXiv) may not represent full landscape of atomistic research; Skill identification bias: LLM-assisted skill extraction and categorization may introduce inconsistencies; Coverage analysis methodology: External LLM (GPT-5.4-mini) used for skill identification, results reviewed by human experts—potential for subjective classification; Case study selection: Six case studies chosen to demonstrate capabilities; unclear if representative of real-world usage patterns; Selection bias in literature sampling (500 papers from specific journals may not represent all computational materials research); Coverage analysis depends on LLM skill identification accuracy (GPT-5.4-mini used for decomposition); Funding bias: Shell Inc. and other corporate sponsors funded portions of the research; Selection bias in literature mining: 500 papers sampled from specific journals may not represent all computational materials science research; Coverage analysis dependent on LLM-based skill identification (GPT-5.4-mini), which may introduce hallucination or misclassification; Funding bias: Research supported by Shell Inc., which may influence scope or validation choices; Validation limited to post-hoc literature comparison rather than prospective user studies

Limitations

  • The framework has notable limitations: "Notably, there are around 12% of journal articles with 0 coverage
  • These missing coverages mainly correspond to articles with novel algorithmic developments, experimental methods, and special domains that require less frequent computational skills." Additionally, the authors note that "the simulated conductivities serve strictly as literature-standard computational screening metrics rather than rigorously converged values targeted for direct experimental validation."

Open questions raised

  • The paper identifies that "an open, systematic, and extensible skill framework dedicated to atomistic materials science and computational chemistry remains a critical open challenge" prior to this work. Future directions implicitly include extending coverage to novel algorithmic developments, experimental methods, and special domains with less frequent computational skills (currently 12% coverage gap). The work also suggests extending framework capabilities to additional scientific domains and improving harmonization with evolving general-purpose coding agents.
  • The authors identify that "an open, systematic, and extensible skill framework dedicated to atomistic materials science and computational chemistry remains a critical open challenge." They note that existing agentic infrastructure has bottlenecks: "Most scientific agents were built as monolithic, standalone platforms atop orchestration frameworks such as LangChain, consequently, developers of these platforms are forced to reinvent fundamental software infrastructure." The gap addressed is bridging deterministic reliability of traditional automation with flexible agentic reasoning.
  • The paper identifies that "an open, systematic, and extensible skill framework dedicated to atomistic materials science and computational chemistry remains a critical open challenge." Additional gaps include: (1) lack of harness engineering approaches layered on general-purpose agents rather than monolithic platforms; (2) limited skill libraries for computational sciences beyond bioinformatics; (3) difficulty scaling LLM agents to manage rigor and complexity of atomistic research.
Data: Materials Project database (accessed via mat-db-mp skill); NIST JANAF database (accessed via mat-db-nist-janaf skill); OPTIMADE databases (accessed via mat-db-optimade skill); QMOF database (accessed via chem-db-mof skill); ARC-MOF database (accessed via chem-db-mof skill); ChEMBL database (accessed via drug-db-chembl skill); PubChem database (accessed via drug-db-pubchem skill); PDB database (accessed via drug-db-pdb skill); DUD-E CDK2 benchmark dataset (474 actives and 27,850 property-matched decoys); Li10GeP2S12 potential energy surface dataset; AtomisticSkills source code: https://github.com/learningmatter-mit/AtomisticSkills; Materials Project database (referenced for materials data); QMOF database (metal-organic frameworks); ARC-MOF database; DUD-E CDK2 benchmark (474 actives and 27,850 property-matched decoys); PDB: 1H1S (CDK2 protein structure); Materials Project (referenced but not created in this work); QMOF and ARC-MOF frameworks (external databases used in MOF screening case study); DUD-E CDK2 benchmark (474 actives and 27,850 property-matched decoys, external dataset); Li10GeP2S12 dataset (generated through workflow, details in case study)Code: https://github.com/learningmatter-mit/AtomisticSkillsExtracted from: pdfAgreement 65%

Explore related topics

Related papers