12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Collaborative and Autonomous AI for Science and Innovation: Practices, Challenges, and Future Directions

Xiaoyu Xiong, Hao Wang, Keming Wu, Zhenfei Yang, Hanjie Zhao, Hao Liu et al. · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
2/4
Quality (LMQS)
I
Evidence
0
Citations
0.00
FWCI

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.36227/techrxiv.177155935.57684125/v1

Methodology & findings

Study design

Narrative literature review and systematic survey of AI applications across stages of the scientific pipeline (knowledge acquisition, hypothesis generation, experiment design, scientific communication).

Main result

The survey identifies two complementary paradigms in AI-driven science: "collaborative AI approaches, where AI systems and human researchers interact in complementary ways across different stages of the scientific workflow, have received comparatively less attention to date. Such approaches can take many forms, including interactive hypothesis refinement, iterative experiment design, guided data analysis, collaborative literature synthesis, and assisted scientific communication. By combining AI's computational capabilities with human contextual judgment, domain expertise, and creative insight, these systems offer opportunities to enhance the reliability, interpretability, efficiency, and overall quality of scientific processes, while addressing some of the limitations inherent in autonomous AI paradigms." Additionally, the findings reveal that "autonomous AI systems can operate at different levels, ranging from supporting specific stages of the scientific pipeline, such as knowledge acquisition and problem formulation, idea and hypothesis generation, experimental design and execution, and scientific communication and presentation to enabling end-to-end execution of the entire scientific process."

Reports effect sizes.

Research paradigm

Interpretivist/constructivist (social science perspective on AI-driven science practices and paradigms)

Author conclusions

The authors conclude that "the convergence of human-centered collaboration with scalable automation offers a promising pathway toward an AI-enabled scientific ecosystem that is both rigorous and imaginative, capable of addressing the complexity and openness inherent in modern science." They further state: "Looking forward, the long-term trajectory of AI-enabled science will depend critically on the institutional and communal foundations that support its practice... there is a pressing need to establish standards, governance structures, and open infrastructures that ensure its contributions are trustworthy, reproducible, and equitable." The authors emphasize that "addressing these issues will be crucial for moving beyond acceleration toward genuine transformation, where AI not only augments productivity but also contributes to the creation of new scientific paradigms."

Risk of bias

Language bias: The survey predominantly focuses on English-language literature and AI systems, potentially underrepresenting non-English scientific traditions; Recency bias: The document is dated 2026 and focuses heavily on recent developments, potentially overweighting novel approaches; Institutional bias: May overrepresent AI developments from well-resourced institutions in developed nations; Publication bias: The review synthesizes published/preprint literature, potentially missing unpublished failures or negative results; Selection bias in coverage of AI systems and tools (English-language bias in indexed literature); Potential overemphasis on published success cases vs. failures; Linguistic and geographic bias: dominance of English-language research in surveyed literature; Representation bias: "Approximately 80% of journal content used for LLM training is in English, whereas 95% of the world's population are not native English speakers"; Publication bias: Review focuses on published systems and tools; unsuccessful or null approaches may be underrepresented; Selection bias: Literature selection methodology not formally specified; skew toward English-language publications and well-established venues; AI-centric bias: Framing emphasizes AI-driven approaches; non-AI methods for scientific discovery receive less emphasis; Anglophone bias: Survey acknowledges that approximately 80% of journal content used for LLM training is in English, and the paper itself is in English

Limitations

  • The survey acknowledges several critical limitations: "Despite this progress, each evaluation approach carries trade-offs
  • Human review captures depth and contextual grounding but lacks scalability
  • LLM-based evaluation provides scalability and flexibility but raises concerns of reliability and circularity, especially if the same families of models act as both generators and evaluators
  • Objective benchmarks offer rigor and reproducibility but risk oversimplification, failing to capture the open-ended creativity central to scientific discovery and innovation
  • End-to-end experimental validation, while the gold standard, remains resource-demanding and infeasible for large-scale adoption." Furthermore, "the most formidable challenge facing the ecosystem described in this section is its profound fragmentation
  • A multidimensional landscape of specialized tools, each highly optimized for a specific task within the research lifecycle

Open questions raised

  • Lack of closed-loop autonomous discovery platforms that can independently interpret results and formulate next research directions
  • Absence of systemic frameworks for documenting and sharing failed experiments and null results
  • Need for improved dynamic knowledge integration and updating in foundation models to keep pace with rapidly evolving scientific literature
  • Limitations in causal and logical reasoning capabilities of current LLMs for scientific inference
  • Significant multilinguality imbalances in AI systems, with 80% of LLM training data in English despite 95% of world population being non-native English speakers
  • Constrained multimodal understanding of heterogeneous scientific data in current foundation models
Data: BigPatent [246]; TLDR [34]; SciQA [16] - leverages ORKG with nearly 170,000 resources; SciQAG [279] - fine-grained QA datasets; IdeaBench [89] - dataset linking target papers with cited references; AI Idea Bench 2025 [223] - 3,495 AI papers with associated inspiration papers; LiveIdeaBench [237] - dataset with 1,180 keywords across 18 scientific domains; Enki - open dataset for archaeological feature recognition [236]; Various scientific repositories mentioned: PubMed Central, arXiv, PubMed Central, ScienceDirect, SpringerLink, Dryad, Zenodo, CORE; BigPatent; TLDR (summarization benchmarks); SciQA benchmark; ORKG (Open Research Knowledge Graph, encoding ~170,000 resources across 15,000+ papers); IdeaBench (3,495 AI papers with inspiration sources); AI Idea Bench 2025 (curated dataset for idea generation evaluation); LiveIdeaBench (1,180 keywords across 18 scientific domains); IdeaBench - benchmark for evaluating LLM scientific idea generation; AI Idea Bench 2025 - 3,495 AI papers with associated inspiration papers; LiveIdeaBench - 1,180 keywords spanning 18 scientific domains; SciQA - leverages ORKG with nearly 170,000 resources; BigPatent - dataset for summarization tasks; TLDR - summarization dataset; BoxingGym - environment-based benchmark for experimental reasoningCode: GitHub - mentioned as de facto standard for code sharing and version control; GitLab - mentioned as alternative repository platform; Llama Factory [36] - standardizes fine-tuning of large language models; ChemOS [231] - orchestration platform for autonomous experimentation; ESCALATE [212] - data-centric pipelines for autonomous experimentation; GitHub (referenced as standard for code sharing and version control); GitLab (referenced as standard for code sharing and version control); Llama Factory (GitHub repository for standardizing LLM fine-tuning); Multiple GitHub repositories cited for AI Scientist v2, AIDE-ML, and other systems; GitLab - mentioned as alternative code repository platform; Llama Factory - open-source framework for fine-tuning LLMs (GitHub reference given)Extracted from: pdfAgreement 70%

Explore related topics

Related papers