Exploring the role of large language models in the scientific method: from hypothesis to discovery
Yanbo Zhang, Sumeer Ahmad Khan, A. K. M. Firoj Mahmud, Huck Yang, Alexander Lavin, Michael Levin et al. · npj Artificial Intelligence · 2025
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1038/s44387-025-00019-5
Methodology & findings
Study design
Narrative review synthesizing recent literature and applications of Large Language Models (LLMs) across scientific domains.
Main result
The paper finds that "LLMs have evolved from tools of convenience-performing tasks like summarising literature, generating code, and analysing datasets-to emerging as pivotal aids in hypothesis generation, experimental design, and even process automation." The authors demonstrate that LLMs can assist across multiple stages of the scientific method: observation (annotation, classification), hypothesis generation (novel drug combinations, nanobody design), experimentation (automated chemical synthesis, gene editing), and automation of discovery loops. Foundation models like Evo and ChemBERT show capability for zero-shot function prediction and multimodal generation tasks, with "the functional activity of Evo-generated CRISPR-Cas molecular complexes and IS200 and IS605 transposable systems was experimentally validated."
Research paradigm
Interpretivist/Critical-realist perspective on AI capabilities and epistemology of scientific discovery
Author conclusions
The authors conclude that "Large Language Models (LLMs) present two contrasting roles in scientific discovery: accuracy in experimental phases and creativity in hypothesis generation." They state that "LLMs and foundation models may come to represent a synthesis of human and artificial intelligence, each amplifying the strengths of the other. With continued research and ethical vigilance, LLMs have the potential to accelerate and deepen scientific discovery, heralding a new era where AI not only supports but inspires new frontiers in science." However, they note a critical caveat: "the scientific community must also decide how much it leaves to AI to drive science, even when associations with 'reasoning', mostly currently undeserved, are made in exchange for the potential to explore hypothesis and solution regions that might otherwise remain unexplored by human exploration alone."
Risk of bias
Selection bias in literature reviewed (may preferentially highlight successful LLM applications); Potential funding bias from AI/tech industry sources; Author affiliation bias (King's College London, research institute may favor AI optimism); Publication bias toward positive results in AI4Science domain; Potential confirmation bias in narrative review methodology; Training data biases in LLMs that may perpetuate research disparities; LLMs tend to provide homogenized critiques across different papers; Hallucinations may introduce false information into scientific processes; Potential bias towards non-native English speakers in scientific publishing; Selection bias in what research is published versus negative results
Limitations
- The authors state several critical limitations: "While carefully designed prompts can accomplish many tasks, they are not robust and reliable enough for complex tasks requiring multiple steps or non-language computations, nor can they explore autonomously." Regarding reasoning, they note "LLMs exhibit severe defects in logical reasoning and serious limitations with respect to common sense reasoning" and can fail at tasks like the "reversal curse" and simple planning problems
- They also acknowledge that "the self-explanation of LLMs is also questionable
- Their explanations are often inconsistent with their behaviours, and we cannot use their explanations to predict their behaviours in counterfactual scenarios." Furthermore, "current state-of-the-art LLMs often fail at simple planning tasks" and "LLMs are not yet good reviewers" for scientific peer review despite their widespread adoption in review writing.
Open questions raised
- The paper identifies the need for clear evaluation metrics and the requirement for collaborative integration of LLMs with human scientific goals across all steps of the scientific process.
- How to develop LLMs capable of autonomous, open-ended exploration of hypothesis spaces
- Advancing LLM reasoning capabilities for novel scientific tasks beyond pattern matching
- Developing robust formal systems and validation methods for LLM-generated hypotheses
- Creating predictive trustworthiness quantification frameworks for LLM agents
- Understanding the extent to which hallucinations can be leveraged creatively versus mitigated
Explore related topics
Related papers
- Preferred reporting items for systematic reviews and meta-analyses: the PRISMA statementDavid Moher · 2009 · 83,271 citations
- PRISMA Extension for Scoping Reviews (PRISMA-ScR): Checklist and ExplanationAndrea C. Tricco · 2018 · 40,391 citations
- Cochrane Handbook for Systematic Reviews of Interventions2019 · 14,420 citations
- PRISMA 2020 explanation and elaboration: updated guidance and exemplars for reporting systematic reviewsMatthew J. Page · 2021 · 10,956 citations
- Updated methodological guidance for the conduct of scoping reviewsMicah D.J. Peters · 2020 · 6,688 citations
- What Is the Impact of ChatGPT on Education? A Rapid Review of the LiteratureChung Kwan Lo · 2023 · 1,725 citations