12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

A Brief History of AI for Scientific Discovery: Open Research, Metrics, and Autonomous Agents

Preprints.org · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
2/4
Quality (LMQS)
I
Evidence
0
Citations

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.20944/preprints202603.1694.v1

Methodology & findings

Study design

Narrative historical analysis and historiographical interpretation.

Main result

The paper traces the historical development of AI for scientific discovery through five eras, arguing that "the history of AI for scientific discovery is the story of science repeatedly handing its bottlenecks to machines, only to discover that each delegation exposes a harder problem underneath." Key developments include DENDRAL's knowledge engineering approach, the shift from rules to data-driven methods, the preprint revolution making science machine-readable, metrics-driven incentive systems, and the emergence of agentic AI systems. Notably, "The AI Scientist, released by Sakana AI in August 2024, which chained together literature search, hypothesis generation, code writing, experiment execution, and manuscript drafting into a single automated pipeline" generated papers costing "approximately fifteen dollars" but containing "citation errors, shallow experiments, and hallucinated results."

Research paradigm

Interpretivist/historical analysis

Author conclusions

The author concludes that "the future of AI in science will not be decided by whether the systems get better. It will be decided by which inheritance prevails. The expert system respect for domain knowledge, the robot scientist's insistence on closing the loop, the open infrastructure's commitment to access, or the human copilot faith that judgement cannot be delegated." The paper emphasizes that "the genuine transformation will come when the research process itself is reorganised around what AI makes possible—continuously updated knowledge graphs, automated experimental programmes, living systematic reviews, forms of scientific communication" rather than simply automating existing workflows. The author cautions that "the real transformation comes not when machines can do what scientists do but when the research process itself is redesigned around AI capabilities."

Risk of bias

Attribution bias in historical narratives (credentials given to visionaries over builders); Geographic bias (AI systems emerge from Western and East Asian institutions; Global North dominance in career infrastructure); Institutional prestige bias in historical accounts (preference for Stanford/MIT/Cambridge lineages); Publication bias in metrics (Google Scholar gaming, h-index manipulation); Selection bias in system examples (prominent systems may not be representative); Historiographical bias: Selection of events and systems included in narrative; Temporal bias: Focus on recent developments (2020s-2026) may overweight contemporary systems; Geographic bias: Most influential systems described emerge from Western and East Asian institutions (San Francisco, Tokyo, Johns Hopkins, Utrecht); Institutional prestige bias: Author notes that KDD originators from industry labs have lower visibility than Stanford/MIT/Cambridge researchers; Framing bias: The paper's central argument (migrating bottleneck metaphor) may shape interpretation of historical events to fit the narrative; Access bias: Primary focus on open source and well-documented systems rather than proprietary or undocumented work; Author selection bias in historical narrative (what systems are highlighted vs. omitted); Potential Western/Global North perspective bias given focus on Stanford, MIT, Cambridge, Tokyo institutions; Recency bias in evaluation of contemporary systems (2024-2026); Disciplinary bias (physics and biology emphasized; social sciences underrepresented)

Limitations

  • The paper acknowledges several limitations: "Independent evaluations of The AI Scientist found serious weaknesses" and notes that "whether these are engineering bugs fixable with better models, or symptoms of a deeper incapacity, is the field's central question." The author also states that "the field needs something equivalent to what CASP was for protein structure prediction: a blind, structured, community validated benchmark for research agent output" indicating that current evaluation mechanisms are inadequate
  • Additionally, the paper notes that "The practical question is whether this constitutes genuine democratisation or merely its appearance" regarding global access to AI tools, suggesting uncertainty about the distribution and accessibility claims being made.

Open questions raised

  • Can autonomous agents produce genuine scientific novelty? (Independent validation at scale still early)
  • How should AI-generated scientific claims be evaluated? (Need for community-validated benchmarks equivalent to CASP)
  • Can peer review survive mounting pressures from AI and volume?
  • What counts as authorship when agents write papers and generate hypotheses?
  • Can tools prevent the hollowing out of scientific craft and Goodhart's Law catastrophes?
  • Whether automation or augmentation will dominate future research practice
Code: GitHub (multiple open source tools mentioned but no specific URLs provided); MIT, Apache 2 licenses (standard for most academic tools); DeepSeek (open source model released January 2025); Llama (referenced as open source model); Qwen (referenced as open source model); GitHub: ASReview (open source framework for systematic reviews); GitHub: Multiple research agents released under MIT, Apache 2 licenses; GitHub: OpenAlex (fully open data and API); ArXiv: The AI Scientist pipeline (open source, August 2024); GitHub: Agent Laboratory (Johns Hopkins); GitHub: Denario (Cambridge and Flatiron Institute); GitHub: LabClaw (Stanford and Princeton); GitHub (general reference to MIT, Apache 2 licenses for academic tools); LangChain (agentic AI framework); AutoGen/AG2 (agentic AI framework); CrewAI (agentic AI framework); LangGraph (mentioned as used by Denario)Extracted from: pdfAgreement 72%

Explore related topics

Related papers