12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

AI Scientists as Engines of Discovery: A Case for Development within Reformed Institutions

Raúl Jiménez, Boris Bolliet, Francisco Villaescusa-Navarro, Rabih Zbib, Benjamin Wandelt, David N. Spergel et al. · ArXiv.org · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
3/4
Quality (LMQS)
I
Evidence
0
Citations
0.00
FWCI

Methodology & findings

Study design

Conceptual analysis and case study of a prototype system (Denario).

Primary method

Design science approach with multi-agent architecture design; informed by philosophy of science and institutional analysis

Main result

The paper argues that "agentic artificial intelligence (AI) systems are beginning to assist, accelerate, and partially automate scientific discovery, performing tasks that span literature synthesis, code generation, data analysis, hypothesis proposal, and model criticism." The authors contend that "suitably designed multi-agent systems may evolve from passive computational tools into 'AI scientists' that can expand the hypothesis-generating and verification capacity of science," but emphasize that "such systems must be developed and deployed within a scientific ecosystem fit for purpose: institutions must be redesigned for verification, accountability, interpretability, and dual-use safety."

Research paradigm

Critical realism with philosophical grounding in epistemology and ethics of science

Author conclusions

The authors conclude that "the capability is there: AI scientists will be developed. The issue is who will be in the driver's seat and what the resulting scientific enterprise will look like." They state that "the institutions of science must be redesigned in parallel, around verification, accountability, interpretability, dual-use safety, methodological diversity, and the protection of human judgment." Most fundamentally, they assert: "The point is not to choose between human and machine science. The point is to build the conditions under which both can do their best work together."

Risk of bias

The paper is authored by researchers actively developing AI systems for science (Denario), creating potential conflict of interest in promoting AI capabilities. The analysis relies primarily on philosophical argumentation rather than empirical evidence. Selection bias may exist in choice of examples and historical parallels cited to support the thesis.; The authors are primarily cosmologists and computer scientists with potential disciplinary bias toward computational/data-intensive fields; No systematic engagement with social science perspectives on scientific institutions; Speculative future-oriented claims not yet empirically validated; Limited discussion of how institutional incentives may bias adoption

Limitations

  • The authors acknowledge that "this essay itself represents a snapshot in time within this rapidly evolving landscape
  • None of the categories we evoke are eternal or impermeable." Current AI systems "do not yet break" the cognitive bottleneck of science "in isolation: they assist but do not reason, they generate but do not falsify reliably, they interpolate but do not extrapolate." The paper also notes that "what an effective authorization interface looks like in a working lab environment is unresolved."

Open questions raised

  • The paper identifies six major governance challenges requiring concrete institutional solutions: (1) dual-use risk mitigation, (2) autonomous experimentation oversight, (3) publication flooding and hallucinated results, (4) methodological homogenization, (5) knowledge collapse from reduced expertise acquisition, and (6) unsafe self-improvement cycles. It also identifies the need for better understanding of effective human-in-the-loop authorization interfaces and the development of AI-aware peer review mechanisms.
  • Unclear effects of mandatory AI disclosure requirements on researcher behavior
  • Lack of understanding about what effective authorization interfaces should look like in working lab environments
  • Need for open benchmarking of hybrid human-machine peer review systems
  • Underdeveloped frameworks for assessing novelty in AI-generated research
  • Limited guidance on protecting epistemic diversity as AI systems converge on similar methods
Code: Denario framework referenced at https://astropilot-ai.github.io/DenarioPaperPage/; https://astropilot-ai.github.io/DenarioPaperPage/Extracted from: pdfAgreement 60%

Explore related topics

Related papers