12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

EvoSci: A Bio-Inspired Multi-Agent Framework for the Evolution of Scientific Discovery

Xiaoyu Xiong, Yuqi Ren, Deyi Xiong · arXiv (Cornell University) · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
1/4
Quality (LMQS)
D
Evidence
0
Citations
0.00
FWCI

Methodology & findings

Study design

Multi-agent simulation framework with computational experiments across ten research topics, systematic evaluation using expert-simulated peer review and tournament-style ranking, ablation studies on component contributions, and qualitative analysis of idea evolution across iterative rounds..

Primary method

Design science approach with iterative system design, computational experimentation, and comparative evaluation

Main result

EvoSci consistently generates more novel and impactful research ideas than strong baselines. "Equipped with DeepSeek-v3, EvoSci achieves the highest overall peer-review scores (ICLR 4.90 / NeurIPS 3.95), surpassing the next best baseline (4.68 / 3.72) by a large margin, and maintaining consistent advantages in terms of Elo-based ranking metrics (Avg Wins 4.19)." The system achieves 5-10% relative gains in ICLR Overall scores and 10-15% gains in NeurIPS Overall scores compared to baselines.

Research paradigm

Design science / computational experimentation

Author conclusions

"In this study, we have presented EvoSci, a multiagent, feedback-driven, and bio-inspired evolutionary framework for automated scientific discovery. The framework conceptualizes scientific discovery as a problem-oriented process, integrates heterogeneous research agents that emulate real-world laboratory roles, and employs multi-round feedback with evolutionary operations to support continuous and open-ended exploration. Extensive experiments across ten scientific domains show that EvoSci consistently outperforms strong baselines in idea validity, excitement, and overall quality."

Risk of bias

Evaluator bias in LLM-based peer review simulation; Selection bias in choice of ten research topics; Potential bias toward certain backbone LLM models; Evaluation bias: LLM-based reviewers may exhibit systematic biases in assessing novelty and feasibility; Selection bias: Choice of ten research topics may not represent all scientific domains equally; Baseline configuration: Fairness of baseline comparisons depends on proper implementation of baseline methods; Tournament evaluation: Pairwise comparison methodology may not capture all quality dimensions; Anthropic/OpenAI API dependency: Results may reflect API model biases and limitations; Evaluation relies heavily on LLM-based reviewers (GPT-4o, DeepSeek-v3, Qwen3-max), which may have built-in biases; Selection of ten research topics from AI Scientist may not represent broader scientific domains equally; Baseline implementations may not represent optimal versions of comparative systems; Meta-reviewer aggregation mechanism not independently validated against human expert judgment; Tournament-style ranking uses arbitrary pairwise comparison prompts that may favor certain types of ideas

Limitations

  • "However, due to its broad cross-domain exploration, the framework sometimes produces ideas with lower practical feasibility, suggesting a trade-off between creativity and applicability." Additionally, the authors note that "a key challenge in achieving such continuous evolution lies in establishing more objective and high-quality evaluation mechanisms that allow LLM-based agents to better assess their own reasoning and outputs."

Open questions raised

  • Limited objective evaluation mechanisms for LLM-based agent assessment
  • Need for enhanced interdisciplinary knowledge integration through structured knowledge representations
  • Requirement for stronger causal reasoning to increase scientific rigor and interpretability
  • Development of more open-ended iterative mechanisms for long-term autonomous scientific discovery
  • Future work should focus on: (1) enhancing EvoSci's ability to reason and operate across disciplines through improved interdisciplinary knowledge integration; (2) strengthening causal reasoning to increase scientific rigor and interpretability; (3) developing more open-ended iterative mechanisms for long-term autonomous scientific discovery; (4) establishing more objective and high-quality evaluation mechanisms for LLM-based self-assessment.
  • Future work should focus on: (1) enhancing EvoSci's ability to reason and operate across disciplines through improved interdisciplinary knowledge integration, (2) strengthening causal reasoning to increase scientific rigor and interpretability, and (3) developing more open-ended iterative mechanisms for long-term autonomous scientific discovery. The authors specifically identify that "establishing more objective and high-quality evaluation mechanisms that allow LLM-based agents to better assess their own reasoning and outputs" is essential for effective self-improvement.
Data: Digital Scientist dataset (based on AMiner Computer Science dataset with 156 representative scientists); Technical Terminology dataset (Artificial Intelligence Terminology Database); Ten experimental task settings from AI Scientist (Lu et al., 2024); Digital Scientist dataset (from VirSci team, based on AMiner Computer Science dataset with 1,712,433 authors and 2,092,356 papers, filtered to 156 representative scientists); Artificial Intelligence Terminology Database (for technical terminology extraction and analysis)Extracted from: pdfAgreement 67%

Explore related topics

Related papers