AutoSci: A Memory-Centric Agentic System for the Full Scientific Research Lifecycle
Weitong Qian, Beicheng Xu, Zhongao Xie, Bowen Fan, Guozheng Tang, Jiale Chen et al. · ArXiv.org · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
Methodology & findings
Study design
End-to-end case study evaluation combined with automated paper-level review proxy.
Primary method
Design science research with iterative system development, case study validation, and automated evaluation proxy
Main result
AutoSci demonstrates that a memory-centric agentic system can execute complete scientific research lifecycles across multiple domains. The system "generates reviewable paper-level artifacts that receive automated ICLR review scores of 6.3/10 and 5.8/10" for GPU kernel optimization and biomedical drug discovery case studies respectively. The GPU kernel case study achieved "a geometric-mean speedup of 1.52× over matched baselines, or 1.18× after excluding degenerate baselines" with "157/157 exe acc = 1.00 at iter5," while the biomedical case produced a "transparent but incomplete negative-result paper" with useful methodological bounds for future work.
Research paradigm
Design science / Systems engineering
Author conclusions
AutoSci demonstrates that a memory-centric architecture can support full-lifecycle scientific research. "Together, these modules allow the system to conduct research across literature understanding, ideation, experimentation, writing, and submission feedback while preserving reusable knowledge and experience across projects." The authors conclude that "AutoSci can produce reviewable paper-level artifacts from end-to-end research processes," with particular success in structured memory construction, idea evolution tracking, and experimental organization. However, they note that "most remain organized around a single project or a paper-generation pipeline" among prior systems, whereas AutoSci's contribution is enabling persistent cross-project knowledge accumulation and system self-evolution.
Risk of bias
Automated review evaluation using PaperReview.ai as proxy rather than formal peer review; Limited to two case studies in specific domains (GPU optimization and biomedical drug discovery); Use of single LLM model (Claude Opus 4.7) for all agent instantiations; Simulated research submissions rather than real peer review submissions; Single model dependency (Claude Code/Opus 4.7); Limited hardware diversity (single GPU families tested); Automated review proxy may not reflect genuine peer review quality; Case studies selected from authors' own research domains; No ground truth comparison against human researcher performance; Limited domain coverage: only two case studies in GPU optimization and biomedical drug discovery; Proxy evaluation risk: using automated PaperReview.ai scores as substitute for real peer review; Single model dependency: both case studies use Claude Opus 4.7 exclusively; Hardware constraints: GPU study limited to single hardware family (NVIDIA A40/4060); Benchmark suite limitations: GPU optimization only evaluated on TritonBench-G workload; No comparison to human researcher baselines or alternative agent systems; Negative result case (biomedical) not subjected to full experimental execution across all planned benchmarks
Limitations
- AutoSci has important limitations
- "First, the current implementation is built as a skill package on top of general-purpose coding and reasoning agents
- This makes the system easy to deploy and inspect, but it is not yet a science-specialized agent foundation." Additionally, "evaluation remains underdeveloped
- Existing benchmarks do not yet adequately measure the separate capabilities required by a full research system, including literature understanding, ideation, experimentation, and writing
- We therefore rely on end-to-end case studies and automated review as a paper-level proxy." The automated reviews also identified domain-specific weaknesses: the GPU case was limited by "a single model, hardware family, and benchmark suite," while the biomedical case was limited by "one main scorer/readout and deferred comparator experiments."
Open questions raised
- Need for science-specialized agent foundation models rather than general-purpose coding agents
- Underdeveloped evaluation frameworks for measuring capabilities of full research systems
- Lack of adequate benchmarks measuring literature understanding, ideation, experimentation, and writing
- Need to accumulate benchmark tasks from real user workflows for future evaluation
- Development of adequate benchmarks to measure separate capabilities in literature understanding, ideation, experimentation, and writing
- Accumulation of benchmark tasks from real user workflows for comprehensive evaluation
Explore related topics
Related papers
- PRISMA Extension for Scoping Reviews (PRISMA-ScR): Checklist and ExplanationAndrea C. Tricco · 2018 · 40,391 citations
- Rayyan—a web and mobile app for systematic reviewsMourad Ouzzani · 2016 · 24,664 citations
- Cochrane Handbook for Systematic Reviews of Interventions2019 · 14,420 citations
- Guidance for conducting systematic scoping reviewsMicah D.J. Peters · 2015 · 7,472 citations
- Updated methodological guidance for the conduct of scoping reviewsMicah D.J. Peters · 2020 · 6,688 citations
- Which academic search systems are suitable for systematic reviews or meta‐analyses? Evaluating retrieval qualities of Google Scholar, PubMed, and 26 other resourcesMichael Gusenbauer · 2019 · 2,116 citations