12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

PaperClaw: Harnessing Agents for Autonomous Research and Human-in-the-Loop Refinement

Weiwei Ye, Hangchen Liu, D C Li, Renhe Jiang · ArXiv.org · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
1/4
Quality (LMQS)
D
Evidence
0
Citations
0.00
FWCI

Methodology & findings

Study design

System design and implementation with LLM-based evaluation.

Primary method

Design science research with iterative system development, comparative evaluation against baseline systems, and ablation studies

Main result

PAPERCLAW produces strong papers both fully autonomously and with human-in-the-loop refinement. In LLM-judge evaluation, "PAPERCLAW, run in human-in-the-loop mode, attains the highest overall score, pairing clear writing with real benchmarks, multiple baselines, ablations, and bootstrap confidence intervals" with an Overall score of 7.0 compared to 4.8 for The AI Scientist, 4.0 for Agent Laboratory, and 4.3 for AUTORESEARCHCLAW.

Research paradigm

Design science / computational research

Author conclusions

"PAPERCLAW harnesses autonomous agents across the entire research lifecycle. By managing domains and brainstorming ideas across them; by driving an iterative hypothesis loop whose reasoning is supplied by an in-cycle, swappable research assistant; by running and managing real experiments; by preserving and reusing the whole project in a full-lifecycle memory; and by compiling a venue-compliant document that cites only validated work, the system turns 'write me a paper' into a structured, resumable, and reusable pipeline."

Risk of bias

LLM judge may have inherent biases in paper assessment; Limited example papers evaluated (small sample size); Papers from GitHub pages may not be representative of typical outputs; System identity was hidden but judge may infer from style or metadata; LLM judge bias - using Claude as evaluator may reflect biases in model training; Sample selection bias - papers chosen from GitHub examples may not be representative; Limited comparison set - competing systems compared against available examples rather than controlled generation; Evaluation metric validity - qualitative scoring anchors (Table 5) lack quantitative grounding; Human-in-the-loop confound - autonomous vs. human-guided runs on different topics, not matched pairs; Evaluation uses LLM judge (Claude) which may have systematic biases in assessing paper quality; Limited comparison set: only 4 baseline systems evaluated; Papers compared are from different topics in autonomous vs. human-in-the-loop ablation, not matched pairs; No evaluation of papers produced on novel/unseen domains; Evaluation limited to example papers available on GitHub pages

Limitations

  • The paper acknowledges that "The remaining gap to a fully trustworthy AI scientist is, fittingly, an empirical one: replacing simulation with execution everywhere, isolating that execution, and evaluating the scientific quality, not just the compliance, of what comes out." Additionally, human-in-the-loop refinement "mainly lifts the weakest autonomous runs rather than the strongest," and the system currently supports "reusable, inspectable knowledge assets rather than learned safeguards" with fully closed-loop learning being "a direction the architecture is designed to support rather than a capability we claim today."

Open questions raised

  • Closing the gap to a fully trustworthy AI scientist through replacing simulation with execution everywhere and isolating that execution
  • Evaluating scientific quality, not just compliance, of autonomous research outputs
  • Implementing fully closed-loop learning where outcomes automatically update policies for hypothesis proposal and experiment design
  • More extensive empirical evaluation of paper quality and scientific validity
  • Need for replacing simulation with execution everywhere
  • Isolating execution in autonomous research systems
Data: Example papers available at: https://github.com/SequenxAI/PaperClaw; Project repository: https://sequenxai.github.io/PaperClawCode: https://github.com/SequenxAI/PaperClaw; https://sequenxai.github.io/PaperClawExtracted from: pdfAgreement 66%

Explore related topics

Related papers