AI Scientists as Engines of Discovery: A Case for Development within Reformed Institutions
Raúl Jiménez, Boris Bolliet, Francisco Villaescusa-Navarro, Rabih Zbib, Benjamin Wandelt, David N. Spergel et al. · ArXiv.org · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
Methodology & findings
Study design
Conceptual analysis and case study of a prototype system (Denario).
Primary method
Design science approach with multi-agent architecture design; informed by philosophy of science and institutional analysis
Main result
The paper argues that "agentic artificial intelligence (AI) systems are beginning to assist, accelerate, and partially automate scientific discovery, performing tasks that span literature synthesis, code generation, data analysis, hypothesis proposal, and model criticism." The authors contend that "suitably designed multi-agent systems may evolve from passive computational tools into 'AI scientists' that can expand the hypothesis-generating and verification capacity of science," but emphasize that "such systems must be developed and deployed within a scientific ecosystem fit for purpose: institutions must be redesigned for verification, accountability, interpretability, and dual-use safety."
Research paradigm
Critical realism with philosophical grounding in epistemology and ethics of science
Author conclusions
The authors conclude that "the capability is there: AI scientists will be developed. The issue is who will be in the driver's seat and what the resulting scientific enterprise will look like." They state that "the institutions of science must be redesigned in parallel, around verification, accountability, interpretability, dual-use safety, methodological diversity, and the protection of human judgment." Most fundamentally, they assert: "The point is not to choose between human and machine science. The point is to build the conditions under which both can do their best work together."
Risk of bias
The paper is authored by researchers actively developing AI systems for science (Denario), creating potential conflict of interest in promoting AI capabilities. The analysis relies primarily on philosophical argumentation rather than empirical evidence. Selection bias may exist in choice of examples and historical parallels cited to support the thesis.; The authors are primarily cosmologists and computer scientists with potential disciplinary bias toward computational/data-intensive fields; No systematic engagement with social science perspectives on scientific institutions; Speculative future-oriented claims not yet empirically validated; Limited discussion of how institutional incentives may bias adoption
Limitations
- The authors acknowledge that "this essay itself represents a snapshot in time within this rapidly evolving landscape
- None of the categories we evoke are eternal or impermeable." Current AI systems "do not yet break" the cognitive bottleneck of science "in isolation: they assist but do not reason, they generate but do not falsify reliably, they interpolate but do not extrapolate." The paper also notes that "what an effective authorization interface looks like in a working lab environment is unresolved."
Open questions raised
- The paper identifies six major governance challenges requiring concrete institutional solutions: (1) dual-use risk mitigation, (2) autonomous experimentation oversight, (3) publication flooding and hallucinated results, (4) methodological homogenization, (5) knowledge collapse from reduced expertise acquisition, and (6) unsafe self-improvement cycles. It also identifies the need for better understanding of effective human-in-the-loop authorization interfaces and the development of AI-aware peer review mechanisms.
- Unclear effects of mandatory AI disclosure requirements on researcher behavior
- Lack of understanding about what effective authorization interfaces should look like in working lab environments
- Need for open benchmarking of hybrid human-machine peer review systems
- Underdeveloped frameworks for assessing novelty in AI-generated research
- Limited guidance on protecting epistemic diversity as AI systems converge on similar methods
Explore related topics
Related papers
- Practical and ethical challenges of large language models in education: A systematic scoping reviewLixiang Yan · 2023 · 699 citations
- ChatGPT in education: Strategies for responsible implementationMohanad Halaweh · 2023 · 576 citations
- Nonhuman “Authors” and Implications for the Integrity of Scientific Publication and Medical KnowledgeAnnette Flanagin · 2023 · 399 citations
- Using AI to write scholarly publicationsMohammad Hosseini · 2023 · 264 citations
- Fighting reviewer fatigue or amplifying bias? Considerations and recommendations for use of ChatGPT and other large language models in scholarly peer reviewMohammad Hosseini · 2023 · 209 citations
- The ethics of disclosing the use of artificial intelligence tools in writing scholarly manuscriptsMohammad Hosseini · 2023 · 202 citations