12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Autonomous chemical research with large language models

Daniil A. Boiko, Robert MacKnight, Ben Kline, Gabriel dos Passos Gomes · Nature · 2023

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
2/4
Quality (LMQS)
E
Evidence
809
Citations
71.26
FWCI
Top 10%
Impact

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1038/s41586-023-06792-0

Methodology & findings

Study design

Multi-method empirical study combining: (1) Benchmark testing of LLM synthesis planning (7 compounds, multiple LLM variants); (2) Documentation retrieval system development and validation using vector embeddings and distance-based search; (3) Robotic liquid handler control experiments with increasing complexity (drawing tasks, color identification); (4) Integrated chemical experimentation with Suzuki–Miyaura and Sonogashira reactions; (5) Optimization game experiments using two published reaction datasets (Suzuki flow synthesis and Buchwald–Hartwig amination) with 20-iteration trials; (6) Computational experiments on large compound libraries.

Sample

> 1000, 4 groups

Primary method

Normalized advantage metric (NA = (yi - mean(y))/(max(y) - mean(y))); normalized maximum advantage (NMA); qualitative analysis of reagent selection patterns; Bayesian optimization baseline comparison; vector-based documentation retrieval with embedding similarity; error bars showing standard deviation for synthesis planning benchmark.

Main result

Coscientist demonstrates versatile autonomous experimental capabilities across six diverse tasks. The study found that "Coscientist showcases its potential for accelerating research across six diverse tasks, including the successful reaction optimization of palladium-catalysed cross-couplings, while exhibiting advanced capabilities for (semi-)autonomous experimental design and execution." Specifically, the GPT-4-powered Web Searcher achieved maximum scores across all trials for acetaminophen, aspirin, nitroaniline and phenolphthalein, and gas chromatography-mass spectrometry analysis confirmed successful formation of target products for both Suzuki and Sonogashira cross-coupling reactions. In optimization experiments, normalized advantage values increased over time, suggesting the model can effectively reuse collected information to guide future actions.

Reports effect sizes.

Research paradigm

Positivist/empiricist; pragmatist engineering approach combining computational methods with physical experimentation

Author conclusions

The authors conclude: "In this paper, we presented a proof of concept for an artificial intelligent agent system capable of (semi-)autonomously designing, planning and multistep executing scientific experiments. Our system demonstrates advanced reasoning and experimental design capabilities, addressing complex scientific problems and generating high-quality code. These capabilities emerge when LLMs gain access to relevant research tools, such as internet and documentation search, coding environments and robotic experimentation platforms. The development of more integrated scientific tools for LLMs has potential to greatly accelerate new discoveries."

Risk of bias

Subjective labeling of synthesis quality (acknowledged by authors); Potential data leakage: unclear if GPT-4 training data contained the Suzuki and Buchwald–Hartwig datasets used for evaluation; Selection of only 7 compounds for synthesis benchmarking (small test set); Limited representativeness of compound space in optimization tasks (5.2% and 6.9% of total reaction space explored); Possible confounding from manual plate movement between automated steps; Funding bias: authors have financial interests (co-founders of aithera.ai; one author on Emerald Cloud Lab advisory board); Potential training data contamination for Suzuki and Buchwald-Hartwig datasets; Subjective labeling of synthesis quality (5-point scale acknowledged as inherently subjective); Limited diversity in experimental tasks (primarily focused on cross-coupling reactions); Small sample sizes for some comparisons (e.g., GPT-3.5 had limited data points due to JSON formatting failures); No blinding of model identity during evaluation; Selection of published datasets may not represent real-world experimental complexity; Training data leakage: Unclear whether GPT-4 training data contains information from reaction datasets used in evaluation, potentially inflating performance metrics; Subjective labeling: Acknowledged by authors that synthesis quality scale labeling is 'inherently subjective' across different evaluators; Limited compound space: Experimental validation limited to small set of available reagents and reactions; Model-specific results: Performance heavily dependent on GPT-4 capabilities; generalization to other LLMs limited (GPT-3.5 and Falcon-40B performed significantly worse); Publication bias potential: Only successful experiments reported; frequency of failures not systematically documented

Limitations

  • The authors note that "although this example requires Coscientist to reason on which reagents are most suitable, our experimental capabilities at that point limited the possible compound space to be explored." Additionally, "Although our setup is not yet fully automated (plates were moved manually), no human decision-making was involved." The paper states that "It is unclear if the GPT-4 training data contain any information from these datasets" when discussing the optimization experiments
  • Regarding safety and data availability, the authors state: "Because of safety concerns, data, code and prompts will be only fully released after the development of US regulations in the field of artificial intelligence and its scientific applications."

Open questions raised

  • Extending Planner's action space to leverage reaction databases such as Reaxys or SciFinder for improved performance in multistep syntheses
  • Advanced prompting strategies (ReAct, Chain of Thought, Tree of Thoughts) for improving accuracy
  • Analyzing system's previous statements as approach to improving accuracy
  • Development of automated quality control techniques in cloud laboratories
  • Need for optimization of experimental parameters (column chemistry, buffer system, gradient)
  • Testing on larger, more diverse compound libraries and reaction spaces
Data: Suzuki reaction dataset from Perera et al. (2018) - flow synthesis with varying ligands, reagents/bases, solvents; Buchwald–Hartwig reaction dataset from Ahneman et al. (2018) - variations in ligands, additives, bases; Emerald Cloud Lab sample catalogue (1,110 model samples); Generated experiment outputs available at https://github.com/gomesgroup/coscientist; Suzuki reaction dataset from Perera et al. (2018) - used for optimization experiments; Buchwald-Hartwig reaction dataset from Doyle et al. (2018) - used for optimization experiments; Generated outputs available at https://github.com/gomesgroup/coscientist; Suzuki reaction dataset (Perera et al., flow synthesis with varying ligands, reagents, bases and solvents); Buchwald-Hartwig reaction dataset (Doyle et al., variations in ligands, additives and bases); ECL cloud laboratory sample catalogue (1,110 Model samples); Results and supplementary outputs: https://github.com/gomesgroup/coscientistCode: https://github.com/gomesgroup/coscientist - Simpler implementation and generated outputs for quantitative analysis; https://github.com/gomesgroup/coscientist (simpler implementation and generated outputs); Full code/prompts to be released after US AI regulations development; https://github.com/gomesgroup/coscientist (simpler implementation and generated outputs for quantitative analysis)Extracted from: pdfAgreement 49%

Explore related topics

Related papers