Agentic AI Scientists Are Not Built For Autonomous Scientific Discovery
Harshit Bisht, Vinay Kumar, Kevin Maik Jablonka, Mausam, N. M. Anoop Krishnan · ArXiv.org · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
Methodology & findings
Study design
Position paper with empirical evidence from a computational experiment (hypothesis hivemind).
Main result
The authors argue that "current agentic systems are not built for autonomous natural science—not because current systems lack scale or tooling, but because the shortcomings are inherent to the training and deployment strategy." They identify four central challenges: "(1) Problem selection is influenced by the McNamara fallacy; (2) Agents are built on large language models (LLMs) whose training corpora omit tacit procedural and failure knowledge of laboratory practice; (3) Preference optimisation during post-training compresses output diversity toward consensus; and (4) Most scientific benchmarks measure single-turn prediction accuracy and lack feedback from physical experiments back to the computational model." The hypothesis hivemind experiment demonstrates that "inter-model similarities remain high despite the desired diversity in task outputs," with frontier models from independent providers converging semantically on both interpretive and open-ended hypothesis generation tasks, with cosine similarities ranging from 0.61 to 0.91 for hypothesis recovery and 0.57 to 0.79 for novel hypothesis generation.
Research paradigm
Critical analysis of AI capabilities and limitations in scientific discovery; epistemologically grounded in philosophy of science (Popperian falsificationism, Polanyian tacit knowledge)
Author conclusions
The authors conclude that "current agentic systems are not built for autonomous natural science—not because current systems lack scale or tooling, but because the shortcomings are inherent to the training and deployment strategy." They argue that AI systems function effectively as co-scientists but not as autonomous agents, and that "until those [design commitments addressing the identified gaps] are in place, the co-scientist model is an accurate description of what the collaboration requires: human judgment that supplements models where they are the most limited." They recommend four concrete interventions: "the use of scientific simulations as verifiers for training, the design of persistent world models that represent the shifting objectives governing real investigations, the establishment of a centralized preregistration repository for all AI-generated hypotheses, and application driven by scientific need rather than tool affordance."
Risk of bias
Selection bias in running example: SSE discovery chosen as illustrative but not systematically justified as representative of all scientific domains; Dataset bias in hypothesis hivemind experiment: limited to 50 papers from single conference track (NeurIPS 2025 AI4Mat), may not generalize to other fields; Model selection bias: only frontier models from two major providers (Anthropic, OpenAI) tested; smaller or specialized models excluded; Embedding model bias: reliance on text-embedding-3-small for semantic comparison may introduce systematic biases in similarity assessment; Publication bias critique is a key argument but not empirically quantified in this paper itself; Confirmation bias risk: authors construct arguments supporting predetermined conclusion that autonomous AI scientists are not viable; counterarguments presented but not equally developed; Selection bias in choice of NeurIPS AI4Mat papers as dataset (narrow domain representation); Embedding model bias: text-embedding-3-small may not capture all semantic distinctions equally; Publication bias in scientific literature that training data reflects; Annotation bias in preference optimization procedures (RLHF/DPO); Confirmation bias in human scientific practice (acknowledged but not fully addressed); Researcher position bias: authors advocate for co-scientist model, may frame evidence accordingly; Dataset bias: Experiment limited to AI4Mat workshop papers, not representative of all scientific domains; Model selection bias: Only two providers (Anthropic and OpenAI) tested; other LLM providers not represented; Task bias: Tasks designed for hypothesis generation may not generalize to other scientific reasoning tasks; Publication bias: Running example (SSE discovery) is from well-studied field with significant published literature, which may exacerbate the discussed consensus convergence; Annotator bias: Post-training via RLHF/DPO encodes preferences of human annotators who are themselves products of scientific consensus; Training data bias: LLMs trained predominantly on published positive results, omitting tacit and failure knowledge
Limitations
- The paper acknowledges but does not exhaustively detail its own scope constraints
- While the authors critique "scale and scaffolding" arguments, they recognize that "improvements in context length, retrieval reliability, and tool-use accuracy will improve capabilities." The hypothesis hivemind experiment, while demonstrating convergence, is limited to 50 papers from a single conference track and 6 models
- the authors do not claim this represents all frontier models or all scientific domains
- The paper relies heavily on conceptual arguments and cited examples rather than systematic quantitative analysis of all claimed phenomena
- The running example of solid-state electrolyte discovery, while illustrative, is not empirically validated against the proposed solutions
- The authors acknowledge legitimate concerns around using peer review data: "We recognise that using such data for training involves legitimate concerns around reviewer consent, privacy, and copyright which must be addressed before such data is used."
Open questions raised
- Methods for capturing tacit knowledge of laboratory practice in training corpora (suggests video records of experimental procedures as a direction)
- Mechanisms for incorporating failure knowledge beyond published literature (requires changes to scientific publishing incentive structures)
- Evaluation frameworks that measure multi-turn scientific reasoning and feedback loops rather than single-turn prediction accuracy
- Design of persistent world models representing shifting objectives in scientific investigations
- Counterfactual reasoning capabilities in scientific agents (noted as currently limited)
- Closing the in silico–in vitro performance gap through mechanistic feedback loops from physical experiments
Explore related topics
Related papers
- What Is the Impact of ChatGPT on Education? A Rapid Review of the LiteratureChung Kwan Lo · 2023 · 1,725 citations
- Artificial intelligence in higher education: the state of the fieldHelen Crompton · 2023 · 1,378 citations
- A comprehensive AI policy education framework for university teaching and learningCecilia Ka Yuk Chan · 2023 · 1,160 citations
- Ethics of AI in Education: Towards a Community-Wide FrameworkW. Holmes · 2021 · 1,056 citations
- The effects of over-reliance on AI dialogue systems on students' cognitive abilities: a systematic reviewChunpeng Zhai · 2024 · 1,009 citations
- Shaping the Future of Education: Exploring the Potential and Consequences of AI and ChatGPT in Educational SettingsSimone Grassini · 2023 · 921 citations