Artificial Intelligence agents for biological research: a survey
Cong Qi, Wenbo Wang, Siqi Jiang, Q. Liu, Xun Song, Zhi Wei · Briefings in Bioinformatics · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1093/bib/bbag075
Methodology & findings
Study design
Structured literature review with systematic search across multiple databases (arXiv, PubMed, bioRxiv, Google Scholar) from January 1, 2023 to September 30, 2025.
Sample
N = 115, 1 group
Primary method
Qualitative synthesis through structured taxonomy; thematic analysis organized across five dimensions (task domains, architectural paradigms, evaluation strategies, interaction modes, and resource integration)
Main result
The survey identifies that "AI agents integrate reasoning, planning, tool invocation, and feedback-driven refinement, enabling more adaptive and interactive forms of biological analysis" across clinical analytics, molecular and drug design, multi-omics analysis, and knowledge discovery. The authors found that "AI agents aim to move beyond static inference toward dynamic reasoning and autonomous experimental workflows" compared to foundation models, with multi-agent systems demonstrating superior performance in complex biological tasks such as single-cell transcriptomics analysis through hierarchical task decomposition and specialized agent coordination.
Reports effect sizes.
Research paradigm
interpretivist/constructivist - systematic synthesis of heterogeneous agentic systems in biological research
Author conclusions
The authors conclude that "current research on biological AI agents demonstrates both progress and limitations" and that while "these systems increasingly exhibit autonomy, interpretability, and domain adaptability, their practical application still faces conceptual and technical barriers." They state that the field needs to move toward "biologically grounded, multi-agent, lightweight, and standardized AI systems that function as collaborative scientific partners, supporting robust, transparent, and widely accessible computational discovery." The authors emphasize that "AI agents are not fully autonomous decision-makers but collaborative tools that are limited by built-in safeguards and clear validation processes."
Risk of bias
Selection bias in literature search - may not capture all relevant studies across diverse databases and disciplinary boundaries; Publication bias - tendency toward positive results in published studies; Heterogeneity in agent definitions - authors note fragmentation in how 'agentic' systems are defined across studies; Domain representation bias - clinical applications are over-represented compared to experimental biology due to data availability and workflow maturity; Selection bias: Search limited to specific databases and date range (Jan 2023-Sept 2025); may miss earlier foundational work or non-indexed publications; Inclusion criteria bias: Manual screening by authors may introduce subjective judgment in determining what constitutes 'agentic characteristics'; Publication bias: Focus on published/preprint studies may exclude negative results or unsuccessful agent implementations; Language bias: Search restricted to English-language databases; Selection bias: Studies were manually screened; no mention of dual-reviewer screening protocol; Publication bias: Only published/preprint studies in major databases included (arXiv, PubMed, bioRxiv, Google Scholar); Language bias: Likely English-language publications only (not explicitly stated); Recency bias: Search limited to 2023-2025 timeframe
Limitations
- The authors note that "development of biological AI agents is still in its nascent stage, with relatively limited and fragmented studies" and that "existing agents are typically designed for single tasks or specific data modalities and are deployed in various contexts." They acknowledge that "in high-risk clinical settings, agents still face substantial heterogeneity in design and application." Additionally, "biological AI agents lack a practical evidence base, gene classifiers contain confounding factors, and there is a potential risk of misuse for dual purposes." The authors further state that agents "struggle in multi-step experimental planning, error diagnosis across heterogeneous data modalities, and coordination of multiple reasoning modes."
Open questions raised
- Lack of standardized evaluation frameworks - heterogeneous assessment strategies across systems
- Data privacy and security concerns in clinical and genomic applications
- Reliability and hallucination mitigation in high-stakes biomedical contexts
- Limited evidence base for agent performance in experimental validation cycles
- Scalability challenges for complex multi-step biological workflows
- Need for domain-specific benchmarks and standardized testing protocols
Explore related topics
Related papers
- Preferred reporting items for systematic reviews and meta-analyses: the PRISMA statementDavid Moher · 2009 · 83,271 citations
- PRISMA Extension for Scoping Reviews (PRISMA-ScR): Checklist and ExplanationAndrea C. Tricco · 2018 · 40,391 citations
- Cochrane Handbook for Systematic Reviews of Interventions2019 · 14,420 citations
- PRISMA 2020 explanation and elaboration: updated guidance and exemplars for reporting systematic reviewsMatthew J. Page · 2021 · 10,956 citations
- Updated methodological guidance for the conduct of scoping reviewsMicah D.J. Peters · 2020 · 6,688 citations
- What Is the Impact of ChatGPT on Education? A Rapid Review of the LiteratureChung Kwan Lo · 2023 · 1,725 citations