A Brief History of AI for Scientific Discovery: Open Research, Metrics, and Autonomous Agents
Preprints.org · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.20944/preprints202603.1694.v1
Methodology & findings
Study design
Narrative historical analysis and historiographical interpretation.
Main result
The paper traces the historical development of AI for scientific discovery through five eras, arguing that "the history of AI for scientific discovery is the story of science repeatedly handing its bottlenecks to machines, only to discover that each delegation exposes a harder problem underneath." Key developments include DENDRAL's knowledge engineering approach, the shift from rules to data-driven methods, the preprint revolution making science machine-readable, metrics-driven incentive systems, and the emergence of agentic AI systems. Notably, "The AI Scientist, released by Sakana AI in August 2024, which chained together literature search, hypothesis generation, code writing, experiment execution, and manuscript drafting into a single automated pipeline" generated papers costing "approximately fifteen dollars" but containing "citation errors, shallow experiments, and hallucinated results."
Research paradigm
Interpretivist/historical analysis
Author conclusions
The author concludes that "the future of AI in science will not be decided by whether the systems get better. It will be decided by which inheritance prevails. The expert system respect for domain knowledge, the robot scientist's insistence on closing the loop, the open infrastructure's commitment to access, or the human copilot faith that judgement cannot be delegated." The paper emphasizes that "the genuine transformation will come when the research process itself is reorganised around what AI makes possible—continuously updated knowledge graphs, automated experimental programmes, living systematic reviews, forms of scientific communication" rather than simply automating existing workflows. The author cautions that "the real transformation comes not when machines can do what scientists do but when the research process itself is redesigned around AI capabilities."
Risk of bias
Attribution bias in historical narratives (credentials given to visionaries over builders); Geographic bias (AI systems emerge from Western and East Asian institutions; Global North dominance in career infrastructure); Institutional prestige bias in historical accounts (preference for Stanford/MIT/Cambridge lineages); Publication bias in metrics (Google Scholar gaming, h-index manipulation); Selection bias in system examples (prominent systems may not be representative); Historiographical bias: Selection of events and systems included in narrative; Temporal bias: Focus on recent developments (2020s-2026) may overweight contemporary systems; Geographic bias: Most influential systems described emerge from Western and East Asian institutions (San Francisco, Tokyo, Johns Hopkins, Utrecht); Institutional prestige bias: Author notes that KDD originators from industry labs have lower visibility than Stanford/MIT/Cambridge researchers; Framing bias: The paper's central argument (migrating bottleneck metaphor) may shape interpretation of historical events to fit the narrative; Access bias: Primary focus on open source and well-documented systems rather than proprietary or undocumented work; Author selection bias in historical narrative (what systems are highlighted vs. omitted); Potential Western/Global North perspective bias given focus on Stanford, MIT, Cambridge, Tokyo institutions; Recency bias in evaluation of contemporary systems (2024-2026); Disciplinary bias (physics and biology emphasized; social sciences underrepresented)
Limitations
- The paper acknowledges several limitations: "Independent evaluations of The AI Scientist found serious weaknesses" and notes that "whether these are engineering bugs fixable with better models, or symptoms of a deeper incapacity, is the field's central question." The author also states that "the field needs something equivalent to what CASP was for protein structure prediction: a blind, structured, community validated benchmark for research agent output" indicating that current evaluation mechanisms are inadequate
- Additionally, the paper notes that "The practical question is whether this constitutes genuine democratisation or merely its appearance" regarding global access to AI tools, suggesting uncertainty about the distribution and accessibility claims being made.
Open questions raised
- Can autonomous agents produce genuine scientific novelty? (Independent validation at scale still early)
- How should AI-generated scientific claims be evaluated? (Need for community-validated benchmarks equivalent to CASP)
- Can peer review survive mounting pressures from AI and volume?
- What counts as authorship when agents write papers and generate hypotheses?
- Can tools prevent the hollowing out of scientific craft and Goodhart's Law catastrophes?
- Whether automation or augmentation will dominate future research practice
Explore related topics
Related papers
- Guidance for conducting systematic scoping reviewsMicah D.J. Peters · 2015 · 7,472 citations
- A SWOT analysis of ChatGPT: Implications for educational practice and researchMohammadreza Farrokhnia · 2023 · 1,171 citations
- ChatGPT and a new academic reality: Artificial Intelligence‐written research papers and the ethics of the large language models in scholarly publishingBrady Lund · 2023 · 769 citations
- Practical and ethical challenges of large language models in education: A systematic scoping reviewLixiang Yan · 2023 · 699 citations
- Unlocking the Power of ChatGPT: A Framework for Applying Generative AI in EducationJiahong Su · 2023 · 550 citations
- Artificial intelligence and the conduct of literature reviewsGerit Wagner · 2021 · 275 citations