12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

ARIS: Autonomous Research via Adversarial Multi-Agent Collaboration

Ruofeng Yang, Yongcan Li, Shuai Li · ArXiv.org · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

10/10
Relevance
1/4
Quality (LMQS)
D
Evidence
0
Citations
0.00
FWCI

Methodology & findings

Study design

Design science research with system architecture, prototype implementation, and observational deployment experience.

Primary method

Design science with iterative refinement; harness engineering

Main result

Aris implements a research harness that coordinates machine-learning research workflows through cross-model adversarial collaboration. The system addresses the central failure mode of autonomous research, which is "plausible unsupported success: a long-running agent can produce claims whose evidential support is incomplete, misreported, or silently inherited from the executor's framing." The report documents "three aspects of Aris: (1) An assurance stack that uses separate executor and reviewer models, including a three-stage process for checking whether claims are supported by evidence; (2) A modular system architecture organized into three layers—execution, orchestration, and assurance—with more than 65 reusable skills; (3) Early deployment experience across three tested executor platforms."

Research paradigm

Design science / engineering pragmatism

Author conclusions

"Aris responds by decomposing the workflow into the three bottlenecks framed in §1—persistent research state, modular execution, and independent assurance—and by adopting a two-role cross-family reviewer-executor pattern as the practical minimum for breaking self-review blind spots." The authors conclude that this conservative design "may understate the capabilities of current agents, but the trade-off favors strictness in a high-rigor field like research: an adversarial reviewer offers a clear quality gain even though adversarial review introduces a harder optimization problem for the executor."

Risk of bias

Lack of controlled evaluation - observational evidence only; No compute-matched baseline comparisons conducted; Potential model-family bias in default executor/reviewer pairings; Confirmation bias in cross-round reviewer context (addressed via fresh-thread option); No comparison to human peer review quality

Limitations

  • "The main limitations are the absence of controlled evaluation and the reliance on observational deployment evidence." The authors note that "Future work includes compute-matched comparisons to estimate the contribution of cross-model heterogeneity (Appendix E), local reviewer models for confidential settings, and user studies of researcher productivity."

Open questions raised

  • Future work includes: (1) compute-matched comparisons to estimate the contribution of cross-model heterogeneity; (2) local reviewer models for confidential settings; (3) user studies of researcher productivity; (4) speculative direction on inserting cross-model accountability primitives between any model output and downstream training-data retention or reward signals; (5) empirical investigation of downstream effects on long-horizon self-improvement.
  • The authors identify several future work directions: (1) compute-matched comparisons to estimate the contribution of cross-model heterogeneity, (2) local reviewer models for confidential settings, (3) user studies of researcher productivity, and (4) adaptation of cross-model accountability primitives (reviewer independence, evidence-to-claim audit) to other domains beyond manuscript review, such as self-improvement mechanisms with explicit oversight layers.
  • Compute-matched comparisons to estimate contribution of cross-model heterogeneity
  • Local reviewer models for confidential/offline settings
  • User studies of researcher productivity
  • Application of cross-model accountability primitives to model self-improvement and training-data retention
Data: No experimental datasets provided; code and documentation available at https://github.com/wanshuiyin/Auto-claude-code-research-in-sleepCode: https://github.com/wanshuiyin/Auto-claude-code-research-in-sleep; https://github.com/ultraworkers/claw-codeExtracted from: pdfAgreement 76%

Explore related topics

Related papers