12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

AutoSci: A Memory-Centric Agentic System for the Full Scientific Research Lifecycle

Weitong Qian, Beicheng Xu, Zhongao Xie, Bowen Fan, Guozheng Tang, Jiale Chen et al. · ArXiv.org · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
1/4
Quality (LMQS)
D
Evidence
0
Citations
0.00
FWCI

Methodology & findings

Study design

End-to-end case study evaluation combined with automated paper-level review proxy.

Primary method

Design science research with iterative system development, case study validation, and automated evaluation proxy

Main result

AutoSci demonstrates that a memory-centric agentic system can execute complete scientific research lifecycles across multiple domains. The system "generates reviewable paper-level artifacts that receive automated ICLR review scores of 6.3/10 and 5.8/10" for GPU kernel optimization and biomedical drug discovery case studies respectively. The GPU kernel case study achieved "a geometric-mean speedup of 1.52× over matched baselines, or 1.18× after excluding degenerate baselines" with "157/157 exe acc = 1.00 at iter5," while the biomedical case produced a "transparent but incomplete negative-result paper" with useful methodological bounds for future work.

Research paradigm

Design science / Systems engineering

Author conclusions

AutoSci demonstrates that a memory-centric architecture can support full-lifecycle scientific research. "Together, these modules allow the system to conduct research across literature understanding, ideation, experimentation, writing, and submission feedback while preserving reusable knowledge and experience across projects." The authors conclude that "AutoSci can produce reviewable paper-level artifacts from end-to-end research processes," with particular success in structured memory construction, idea evolution tracking, and experimental organization. However, they note that "most remain organized around a single project or a paper-generation pipeline" among prior systems, whereas AutoSci's contribution is enabling persistent cross-project knowledge accumulation and system self-evolution.

Risk of bias

Automated review evaluation using PaperReview.ai as proxy rather than formal peer review; Limited to two case studies in specific domains (GPU optimization and biomedical drug discovery); Use of single LLM model (Claude Opus 4.7) for all agent instantiations; Simulated research submissions rather than real peer review submissions; Single model dependency (Claude Code/Opus 4.7); Limited hardware diversity (single GPU families tested); Automated review proxy may not reflect genuine peer review quality; Case studies selected from authors' own research domains; No ground truth comparison against human researcher performance; Limited domain coverage: only two case studies in GPU optimization and biomedical drug discovery; Proxy evaluation risk: using automated PaperReview.ai scores as substitute for real peer review; Single model dependency: both case studies use Claude Opus 4.7 exclusively; Hardware constraints: GPU study limited to single hardware family (NVIDIA A40/4060); Benchmark suite limitations: GPU optimization only evaluated on TritonBench-G workload; No comparison to human researcher baselines or alternative agent systems; Negative result case (biomedical) not subjected to full experimental execution across all planned benchmarks

Limitations

  • AutoSci has important limitations
  • "First, the current implementation is built as a skill package on top of general-purpose coding and reasoning agents
  • This makes the system easy to deploy and inspect, but it is not yet a science-specialized agent foundation." Additionally, "evaluation remains underdeveloped
  • Existing benchmarks do not yet adequately measure the separate capabilities required by a full research system, including literature understanding, ideation, experimentation, and writing
  • We therefore rely on end-to-end case studies and automated review as a paper-level proxy." The automated reviews also identified domain-specific weaknesses: the GPU case was limited by "a single model, hardware family, and benchmark suite," while the biomedical case was limited by "one main scorer/readout and deferred comparator experiments."

Open questions raised

  • Need for science-specialized agent foundation models rather than general-purpose coding agents
  • Underdeveloped evaluation frameworks for measuring capabilities of full research systems
  • Lack of adequate benchmarks measuring literature understanding, ideation, experimentation, and writing
  • Need to accumulate benchmark tasks from real user workflows for future evaluation
  • Development of adequate benchmarks to measure separate capabilities in literature understanding, ideation, experimentation, and writing
  • Accumulation of benchmark tasks from real user workflows for comprehensive evaluation
Data: TritonBench workspace (GPU kernel optimization case); DeepTernary v1.0.0 (biomedical drug discovery case); PROTAC-STAN inference repositories (biomedical drug discovery case); 22-complex unbound PROTAC test set (biomedical drug discovery case); TritonBench workspace (GPU kernel optimization); DeepTernary v1.0.0 and PROTAC-STAN inference repositories (biomedical case); DeepTernary v1.0.0 and PROTAC-STAN inference repositories (biomedical drug discovery); 22-complex unbound PROTAC test set (biomedical)Code: TritonBench workspace (referenced for GPU kernel optimization); DeepTernary v1.0.0 and PROTAC-STAN inference repositories (referenced for biomedical case); Triton 3.2.0 and PyTorch 2.6.0+cu124 (execution environments); TritonBench workspace (mentioned but no URL provided); DeepTernary v1.0.0 (external repository reference); PROTAC-STAN (external repository reference); TritonBench repository (GPU kernel optimization experiments); DeepTernary v1.0.0 repository (biomedical drug discovery); PROTAC-STAN inference repository (biomedical drug discovery)Extracted from: pdfAgreement 64%

Explore related topics

Related papers