AI, agentic models and lab automation for scientific discovery — the beginning of scAInce
Thomas Hartung · Frontiers in Artificial Intelligence · 2025
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.3389/frai.2025.1649155
Methodology & findings
Study design
Narrative review of literature and case studies.
Primary method
No original statistical analyses conducted (narrative review). The paper references multiple statistical and machine learning methods used in cited studies: sparse canonical correlation analysis (sCCA), deep neural networks, graph neural networks, Bayesian optimization, reinforcement learning from human feedback, causal inference methods, active learning algorithms, and Bayesian experimental design.
Main result
The paper finds that "AI systems also promote a more interconnected scientific community. By breaking down barriers related to the accessibility of knowledge, they enable a democratization of information, allowing for a broader base of scientists to engage with cutting-edge research regardless of geographical or institutional boundaries." Additionally, "researchers can devote more time to experimental design and data interpretation rather than sifting through literature, which is accelerating the pace of scientific discovery." The review demonstrates that autonomous systems are reshaping scientific practice, from literature synthesis to experimental execution, with examples including "Insilico Medicine's generatively designed anti-fibrotic ISM001-055 advanced to Phase II trials in 2024-the first AI-invented small molecule to reach that milestone."
Reports confidence intervals.
Research paradigm
Post-positivist; advocates for machine-readable science and algorithmic optimization of research processes
Author conclusions
The author concludes that science is undergoing a qualitative transformation: "We stand on the brink of a qualitative shift in how knowledge is generated. The distinction between reading science and doing science is blurring as agentic AI systems orchestrate both literature review and laboratory execution." The author emphasizes both opportunity and responsibility: "The critical question is not whether science will accelerate but whose science will accelerate and under what safeguards. As editors we have the privilege and the duty to shape this trajectory." The author also advocates for treating scAInce (science optimized for artificial intelligence) as a paradigm shift requiring institutional transformation toward "machine-readable publications as the default scholarly unit" and governance frameworks ensuring "equity-of-access and diversity-of-approach metrics."
Risk of bias
Algorithmic bias in AI-driven literature screening and hypothesis generation; Data bias in training datasets skewing suggestions toward well-studied pathways; Language bias potentially amplifying English-language journal dominance; Selection bias in AI systems trained on incomplete or domain-biased corpora; Potential bias toward machine-tractable research problems over curiosity-driven exploration; Selection bias in choice of exemplar projects featured (favorable cases may be overrepresented); Language bias (English-language publications overrepresented in AI training data); Institutional bias (examples predominantly from well-resourced organizations and corporations); Algorithmic bias in AI systems trained on biased corpora; Publication bias in cited literature (positive results more likely published); Representativeness bias toward data-rich and computationally tractable domains; Algorithmic bias in LLM training data leading to skewed hypotheses; Selection bias in AI-assisted literature screening without human oversight; Language bias favoring English-language publications in algorithmically optimized systems; Geographic bias marginalizing Global South scholarship; Data monopoly concentration in commercial AI platforms; Training data incompleteness and domain bias in models; Potential bias toward machine-tractable problems in research agenda setting
Limitations
- The author identifies several critical limitations: "However, relying on AI for hypothesis generation also poses challenges related to their validation
- Ensuring that these hypotheses are not only innovative but also scientifically valid and testable is crucial
- There is also the need to address potential biases in the data used by LLMs, which could lead to skewed or unrepresentative hypotheses, impacting the direction and integrity of research efforts." Additionally, regarding systematic reviews: "Over-reliance on AI-assisted screening without adequate human oversight risks propagating selection bias or misclassification of studies, particularly when models are trained on incomplete or domain-biased corpora." The paper also warns that "algorithmic optimization may amplify the imbalance, marginalizing scholars from the Global South whose work remains under-represented in major repositories."
Open questions raised
- Need for robust validation frameworks ensuring AI-generated hypotheses are scientifically valid and testable
- Addressing potential biases in LLM training data and avoiding skewed research directions
- Establishing transparency and interpretability standards for AI-driven experimental design decisions
- Developing causal inference tools coupled with generative modeling to move beyond latent-space proximity to genuine causation
- Creating interoperability standards between review platforms for uniform parameter adoption
- Scaling autonomous laboratory infrastructure (self-driving labs) beyond well-resourced institutions through shared infrastructure and public compute-credit programs
Explore related topics
Related papers
- What Is the Impact of ChatGPT on Education? A Rapid Review of the LiteratureChung Kwan Lo · 2023 · 1,725 citations
- Artificial intelligence in higher education: the state of the fieldHelen Crompton · 2023 · 1,378 citations
- Ethics of AI in Education: Towards a Community-Wide FrameworkW. Holmes · 2021 · 1,056 citations
- The effects of over-reliance on AI dialogue systems on students' cognitive abilities: a systematic reviewChunpeng Zhai · 2024 · 1,009 citations
- Shaping the Future of Education: Exploring the Potential and Consequences of AI and ChatGPT in Educational SettingsSimone Grassini · 2023 · 921 citations
- Revolutionizing education with AI: Exploring the transformative potential of ChatGPTTufan Adıgüzel · 2023 · 858 citations