12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Accelerating scientific discovery with Co-Scientist

Juraj Gottweis, Wei-Hung Weng, Alexander Daryin, Tao Tu, Petar Sirkovic, Anatoly Myaskovsky et al. · Nature · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
4/4
Quality (LMQS)
E
Evidence
9
Citations
144.50
FWCI
Top 10%
Impact

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1038/s41586-026-10644-y

Methodology & findings

Study design

Multi-method validation study combining: (1) computational benchmarking using Elo-based tournament evaluation across 203 research goals; (2) expert evaluation on 15 curated biomedical research goals; (3) in vitro biological validation using cancer cell lines (5 AML cell lines and 1 control); (4) human hepatic organoid studies for liver fibrosis; (5) computational hypothesis generation for antimicrobial resistance mechanisms.

Sample

N = 203, 14 groups

Primary method

Elo-based tournament ranking system (pairwise comparisons with multi-turn scientific debates for top hypotheses, single-turn for lower-ranked), non-linear regression curve fitting for dose-response IC50 estimation, Chou-Talalay combination index method for doublet drug interactions, Highest Single Agent (HSA) and Bliss independence models for triplet combinations, Elo auto-evaluation metric. Software: Python 3.11.7, pandas 2.1.4, numpy 1.26.4, seaborn 0.12.2, matplotlib 3.8.0, GraphPad Prism 10.6.0, Julius AI statistical software (accessed November 2025).

Main result

Co-Scientist successfully generated novel, testable hypotheses across three biomedical domains. In drug repurposing for AML, "Binimetinib, which is already approved for the treatment of metastatic melanoma, exhibited an half-maximal inhibitory concentration (IC50) as low as 2 nM in all AML cell lines (except NOMO-1)". For liver fibrosis, the system "successfully identified three novel epigenetic modifiers and drugs targeting them, and two of them exhibited significant anti-fibrotic activity in the hepatic organoids without causing cellular toxicity". For antimicrobial resistance, "Co-Scientist independently and accurately proposed the groundbreaking, top-ranked hypothesis that cf-PICIs interact with diverse phage tails to expand their host range", which "precisely matched the primary discovery of an independent, co-timed genomic and experimental study prior to completing peer-review".

Reports effect sizes.

Research paradigm

Empirical-computational; mixed-methods (computational evaluation + in vitro validation + expert assessment)

Author conclusions

"Co-Scientist represents a promising step towards AI-assisted augmentation of scientists and acceleration of scientific discovery. Its ability to think scientifically, generate novel testable hypotheses across diverse scientific and biomedical domains, some supported by experimental findings, along with the capacity for recursive self-improvement with increasing compute, demonstrates the promise of meaningfully accelerating scientists' endeavors to resolve grand challenges in human health, medicine and science." The authors emphasize this is "a promising step" requiring continued development, robust verification methods, and rigorous peer review integration.

Risk of bias

Publication bias: system limited to open-access literature, missing paywalled studies and negative results; Literature quality bias: system depends on source literature quality which may be 'mixed and contradictory'; Hallucination risk: authors note 'imperfect factuality and the potential for hallucinations' in underlying LLMs; Research direction homogenization: potential for AI to create bias in scientific directions rather than augment; Expert selection bias: experts who curated research goals may have biased initial hypotheses; Limited validation scope: only 3 biomedical applications validated; generalizability unclear; Expert selection bias: Only 7 experts curated the 15 research goals; limited demographic diversity of experts; Publication bias risk: System relies on open-access literature, systematically excluding paywalled research; Evaluation bias: Elo rating is an auto-evaluation metric, not independent ground truth; Confirmation bias: Researchers may have selectively reported successful validations; Limited negative results: Three biomedical applications chosen may represent optimal domains for AI performance; Small sample size for expert evaluation: Only 11 of 15 research goals were assessed by human experts; Model dependency: System built on Gemini; generalization to other LLMs not empirically validated; Limited to open-access literature, introducing publication bias and paywalled research exclusion; Reliance on mixed-quality source literature with potential propagation of erroneous findings; Expert curation of research goals may introduce selection bias in problem selection; Small scale of expert evaluations (n=11 for preference ranking) limits generalizability; In vitro validation may not translate to in vivo efficacy; Potential for AI-generated homogenization of research directions; Elo rating is auto-evaluated metric, not independent ground truth; Model-specific biases from Gemini LLM architecture

Limitations

  • The authors acknowledge that "Co-Scientist's knowledge is constrained by its reliance on open-access scientific literature, which may lead to the omission of critical prior art behind paywalls and a systemic lack of access to negative experimental results." They further state "the quality of generated hypotheses relies on the mixed and contradictory quality of the source literature
  • thus, there is a risk of propagating erroneous or irreproducible findings." Additionally, "Co-Scientist also inherits the intrinsic limitations of its underlying models, including imperfect factuality and the potential for hallucinations." The authors note that "the validation of Co-Scientist's hypotheses, while successful, remains preliminary" and emphasize that "translating these predictions from Co-Scientist into clinical practice will be highly challenging, as the complexity of a disease model, patient heterogeneity, and disease variability cannot be fully captured in such limited in vitro experiments."

Open questions raised

  • Authors identify future directions: (1) Immediate improvements: enhance robustness, literature search breadth, fact-checking against external databases, citation recall; (2) Core capability expansion: integrate agents for reasoning over public databases and multimodal data, implement bioinformatics/data science tasks, use reinforcement learning from human and experimental feedback; (3) Expanded evaluation: assess generalizability across wider scientific disciplines, develop objective automated evaluation metrics beyond ranking systems, engage larger expert cohorts; (4) Integration with lab automation for closed-loop autonomous hypothesis generation and experimental validation.
  • Development of agents with enhanced provenance capabilities to trace claims to specific figures or data within sources
  • Improving reasoning capabilities to address imperfect factuality and hallucinations
  • Developing robust verification methods and rigorous peer review processes for AI-generated hypotheses
  • Expanding evaluations to assess generalizability across wider range of scientific disciplines
  • Developing more objective and automated evaluation metrics beyond current ranking systems
Data: No datasets explicitly stated as publicly available. Authors note: 'The full source code for the Co-Scientist system is not publicly available.' However, 'we are initiating an experimental access program' for scientists interested in accessing Co-Scientist.; No public datasets explicitly provided. The paper states that "to enable and accelerate research on important scientific problems, we are initiating an experimental access program" and that "the full source code for the Co-Scientist system is not publicly available." Access is available via Google APIs and bespoke interfaces over time. The Gemini foundational LLM is publicly available via APIs, Google AI studio, and the Gemini App.Code: Full source code not publicly available. System implemented in Python 3.11.7. Authors provide: comprehensive pseudocode (Supplementary Note 8), exact system prompts (Supplementary Note 9), and note that 'The foundational LLM model, Gemini, used in Co-Scientist is publicly available via APIs, Google AI studio, the Gemini App and other surfaces.'; Full source code for Co-Scientist is not publicly released. Comprehensive pseudocode is provided in Supplementary Note 8. Exact system prompts are provided in Supplementary Note 9. The analysis codes were developed using Python 3.11.7 but are not stated to be available. Gemini model is publicly accessible.; Full source code not publicly available. Authors initiated experimental access program. Comprehensive pseudocode provided in Supplementary Note 8; system prompts provided in Supplementary Note 9. Foundational LLM model (Gemini) publicly available via APIs and Google AI Studio.Extracted from: pdfAgreement 45%

Explore related topics

Related papers