Accelerating scientific discovery with Co-Scientist
Juraj Gottweis, Wei-Hung Weng, Alexander Daryin, Tao Tu, Petar Sirkovic, Anatoly Myaskovsky et al. · Nature · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1038/s41586-026-10644-y
Methodology & findings
Study design
Multi-method validation study combining: (1) computational benchmarking using Elo-based tournament evaluation across 203 research goals; (2) expert evaluation on 15 curated biomedical research goals; (3) in vitro biological validation using cancer cell lines (5 AML cell lines and 1 control); (4) human hepatic organoid studies for liver fibrosis; (5) computational hypothesis generation for antimicrobial resistance mechanisms.
Sample
N = 203, 14 groups
Primary method
Elo-based tournament ranking system (pairwise comparisons with multi-turn scientific debates for top hypotheses, single-turn for lower-ranked), non-linear regression curve fitting for dose-response IC50 estimation, Chou-Talalay combination index method for doublet drug interactions, Highest Single Agent (HSA) and Bliss independence models for triplet combinations, Elo auto-evaluation metric. Software: Python 3.11.7, pandas 2.1.4, numpy 1.26.4, seaborn 0.12.2, matplotlib 3.8.0, GraphPad Prism 10.6.0, Julius AI statistical software (accessed November 2025).
Main result
Co-Scientist successfully generated novel, testable hypotheses across three biomedical domains. In drug repurposing for AML, "Binimetinib, which is already approved for the treatment of metastatic melanoma, exhibited an half-maximal inhibitory concentration (IC50) as low as 2 nM in all AML cell lines (except NOMO-1)". For liver fibrosis, the system "successfully identified three novel epigenetic modifiers and drugs targeting them, and two of them exhibited significant anti-fibrotic activity in the hepatic organoids without causing cellular toxicity". For antimicrobial resistance, "Co-Scientist independently and accurately proposed the groundbreaking, top-ranked hypothesis that cf-PICIs interact with diverse phage tails to expand their host range", which "precisely matched the primary discovery of an independent, co-timed genomic and experimental study prior to completing peer-review".
Reports effect sizes.
Research paradigm
Empirical-computational; mixed-methods (computational evaluation + in vitro validation + expert assessment)
Author conclusions
"Co-Scientist represents a promising step towards AI-assisted augmentation of scientists and acceleration of scientific discovery. Its ability to think scientifically, generate novel testable hypotheses across diverse scientific and biomedical domains, some supported by experimental findings, along with the capacity for recursive self-improvement with increasing compute, demonstrates the promise of meaningfully accelerating scientists' endeavors to resolve grand challenges in human health, medicine and science." The authors emphasize this is "a promising step" requiring continued development, robust verification methods, and rigorous peer review integration.
Risk of bias
Publication bias: system limited to open-access literature, missing paywalled studies and negative results; Literature quality bias: system depends on source literature quality which may be 'mixed and contradictory'; Hallucination risk: authors note 'imperfect factuality and the potential for hallucinations' in underlying LLMs; Research direction homogenization: potential for AI to create bias in scientific directions rather than augment; Expert selection bias: experts who curated research goals may have biased initial hypotheses; Limited validation scope: only 3 biomedical applications validated; generalizability unclear; Expert selection bias: Only 7 experts curated the 15 research goals; limited demographic diversity of experts; Publication bias risk: System relies on open-access literature, systematically excluding paywalled research; Evaluation bias: Elo rating is an auto-evaluation metric, not independent ground truth; Confirmation bias: Researchers may have selectively reported successful validations; Limited negative results: Three biomedical applications chosen may represent optimal domains for AI performance; Small sample size for expert evaluation: Only 11 of 15 research goals were assessed by human experts; Model dependency: System built on Gemini; generalization to other LLMs not empirically validated; Limited to open-access literature, introducing publication bias and paywalled research exclusion; Reliance on mixed-quality source literature with potential propagation of erroneous findings; Expert curation of research goals may introduce selection bias in problem selection; Small scale of expert evaluations (n=11 for preference ranking) limits generalizability; In vitro validation may not translate to in vivo efficacy; Potential for AI-generated homogenization of research directions; Elo rating is auto-evaluated metric, not independent ground truth; Model-specific biases from Gemini LLM architecture
Limitations
- The authors acknowledge that "Co-Scientist's knowledge is constrained by its reliance on open-access scientific literature, which may lead to the omission of critical prior art behind paywalls and a systemic lack of access to negative experimental results." They further state "the quality of generated hypotheses relies on the mixed and contradictory quality of the source literature
- thus, there is a risk of propagating erroneous or irreproducible findings." Additionally, "Co-Scientist also inherits the intrinsic limitations of its underlying models, including imperfect factuality and the potential for hallucinations." The authors note that "the validation of Co-Scientist's hypotheses, while successful, remains preliminary" and emphasize that "translating these predictions from Co-Scientist into clinical practice will be highly challenging, as the complexity of a disease model, patient heterogeneity, and disease variability cannot be fully captured in such limited in vitro experiments."
Open questions raised
- Authors identify future directions: (1) Immediate improvements: enhance robustness, literature search breadth, fact-checking against external databases, citation recall; (2) Core capability expansion: integrate agents for reasoning over public databases and multimodal data, implement bioinformatics/data science tasks, use reinforcement learning from human and experimental feedback; (3) Expanded evaluation: assess generalizability across wider scientific disciplines, develop objective automated evaluation metrics beyond ranking systems, engage larger expert cohorts; (4) Integration with lab automation for closed-loop autonomous hypothesis generation and experimental validation.
- Development of agents with enhanced provenance capabilities to trace claims to specific figures or data within sources
- Improving reasoning capabilities to address imperfect factuality and hallucinations
- Developing robust verification methods and rigorous peer review processes for AI-generated hypotheses
- Expanding evaluations to assess generalizability across wider range of scientific disciplines
- Developing more objective and automated evaluation metrics beyond current ranking systems
Explore related topics
Related papers
- The effects of over-reliance on AI dialogue systems on students' cognitive abilities: a systematic reviewChunpeng Zhai · 2024 · 1,009 citations
- Autonomous chemical research with large language modelsDaniil A. Boiko · 2023 · 809 citations
- Comparing scientific abstracts generated by ChatGPT to real abstracts with detectors and blinded human reviewersCatherine A. Gao · 2023 · 657 citations
- Teacher support and student motivation to learn with Artificial Intelligence (AI) based chatbotThomas K. F. Chiu · 2023 · 617 citations
- Beware of metacognitive laziness: Effects of generative artificial intelligence on learning motivation, processes, and performanceYizhou Fan · 2024 · 419 citations
- What ChatGPT means for universities: Perceptions of scholars and studentsMehmet Fırat · 2023 · 405 citations