12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

A Survey of AI Scientists

Guiyao Tie, Pan Zhou, Lichao Sun · arXiv (Cornell University) · 2025

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
2/4
Quality (LMQS)
I
Evidence
0
Citations

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.48550/arxiv.2510.23045

Methodology & findings

Study design

Narrative scoping review with systematic mapping of literature (2022-2025) onto a structured six-stage methodological framework.

Sample

not–applicable

Main result

The survey identifies "a clear three-phase evolutionary trajectory" of AI Scientist systems development: "from an initial phase of Foundational Modules, focused on task-specific automation, through a period of Closed-Loop Integration, to the current frontier of Scalability, Impact, and Collaboration." The authors establish "a unified, six-stage methodological framework that deconstructs the scientific process into: Literature Review, Idea Generation, Experimental Preparation, Experimental Execution, Scientific Writing, and Paper Generation." Systems such as "The AI Scientist v1 and v2 have demonstrated this capacity by generating entire scientific papers, inclusive of figures, experiments, and internal reviews, with minimal human oversight."

Reports effect sizes.

Research paradigm

Systems and computational epistemology; philosophy of science applied to AI-automated research workflows

Author conclusions

"This survey provides a critical roadmap for the field, intended to guide the next generation of systems toward becoming trustworthy, verifiable, and indispensable partners in human scientific inquiry." The authors conclude that the field has progressed through three phases and identify "dual research thrusts toward both greater machine autonomy and more sophisticated human-in-the-loop synergy." They emphasize that "AI is evolving from an instrument of inquiry into a potential originator of scientific knowledge" while cautioning that the field must address "critical open challenges in robustness, generalizability, and ethical governance."

Risk of bias

Selection bias: Review limited to 2022-2025 timeframe, potentially excluding earlier foundational work; Language bias: Appears to focus on English-language publications from major repositories (arXiv, PubMed); Publication bias: Likely skewed toward published/preprinted systems over failed or unreported attempts; Domain representation bias: May over-represent well-funded domains (chemistry, materials science) with robotic infrastructure; Temporal scope limitation (2022-2025) may exclude foundational work from earlier periods; Selection bias in choosing 'seminal works' for inclusion in the taxonomy; Potential publication bias toward successful/mature systems over unsuccessful attempts; Language bias (predominantly English-language sources from arXiv, PubMed, Semantic Scholar); Domain representation bias (chemistry and materials science noted as most mature; other domains may be underrepresented); Selection bias: The review focuses only on research from 2022-2025, potentially missing earlier foundational work in automated science. Publication bias: Focus on published/preprint works may exclude grey literature or failed systems. Author expertise bias: Categorization of systems into the proposed framework relies on author interpretation of paper contributions rather than independent verification.

Open questions raised

  • Lack of unified taxonomy linking scientific tasks, AI capabilities, agentic systems, and evaluation protocols
  • Need for comprehensive frameworks addressing full-cycle scientific autonomy rather than task-specific automation
  • Open challenges in robustness, generalizability, and ethical governance of autonomous research systems
  • Limited work on human-in-the-loop synergy and collaborative AI-human research partnerships
  • Gaps in reproducibility, transparency, and verifiability standards for AI-generated science
  • Underexplored domains beyond chemistry and materials science
Data: S2ORC (S2ORC corpus for scientific paper parsing); SciDocs (domain-specific corpus for scientific IR evaluation); UMLS and MeSH (biomedical ontologies for knowledge graph grounding); S2ORC (large-scale, full-text corpus); SciDocs (scientific corpus for training); UMLS and MeSH (established ontologies referenced); No empirical datasets are generated or analyzed in this survey. However, the paper references multiple benchmark datasets and systems: DS-1000 [37], MLAgentBench [40], IdeaBench [42], ResearchBench [23], EXP-Bench [47], Auto-Bench [22]. The authors provide a GitHub project: https://github.com/Mr-Tieguigui/Survey-for-AI-ScientistCode: https://github.com/Mr-Tieguigui/Survey-for-AI-Scientist (Project GitHub mentioned in abstract); https://github.com/Mr-Tieguigui/Survey-for-AI-Scientist; Project GitHub: https://github.com/Mr-Tieguigui/Survey-for-AI-ScientistExtracted from: pdfAgreement 53%

Explore related topics

Related papers