12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

The Last Human-Written Paper: Agent-Native Research Artifacts

Jiachen Liu, Jiaxin Pei, Jintao Huang, Chenglei Si, Ao Qu, Xiangru Tang et al. · ArXiv.org · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
1/4
Quality (LMQS)
D
Evidence
0
Citations
0.00
FWCI

Methodology & findings

Study design

Multi-layered mixed-method evaluation across three benchmarks (PaperBench, RE-Bench, METR eval-analysis-public dataset).

Primary method

Design Science with iterative refinement. Four-layer structured protocol design (Cognitive, Physical, Exploration Graph, Evidence layers) informed by taxonomy of reproduction-critical information categories from PaperBench rubrics.

Main result

ARA raises question-answering accuracy from 72.4% to 93.7% and reproduction success from 57.4% to 64.4%. On the understanding layer, "ARA outperforms the baseline at every category and every benchmark, with overall accuracy 93.7% vs. 72.4% (+21.3%) on 450 paired outcomes." On reproduction, "Across all 15 papers with complete paired runs (150 subtasks, 1,743 rubric requirements), ARA achieves a difficulty-weighted success rate of 64.4% vs. 57.4% for the baseline." On extension, "preserved failure traces in ARA accelerate progress, but can also constrain a capable agent from stepping outside the prior-run box depending on the agent's capabilities."

Research paradigm

Design science and artifact-centered research with agent-native computing

Author conclusions

The Agent-Native Research Artifact protocol recasts the primary research object from narrative document to agent-executable knowledge package. "Tolerable for human readers, these costs become critical when AI agents must understand, reproduce, and extend published work." The authors conclude that "research is scaling into a massively parallel enterprise where agents fork, extend, and merge each other's work at machine speed, shifting the bottleneck from individual productivity to artifact operability: narrative PDFs, compiled for sequential human reading, cannot be forked, diffed, or merged, but a structured, lossless artifact can, letting research compound like software."

Risk of bias

Selection bias in benchmark choice (ICML 2024 papers and RE-Bench tasks); potential evaluator bias in blinded judging; capability-dependent results (different models show inverted preferences); beat-reference filter applied to RE-Bench but fairness depends on filter correctness.; Selection bias in paper choice (only 15 of 23 PaperBench papers had companion repositories); task difficulty stratification may not be representative; baseline construction (LLM-synthesized paper writeups for RE-Bench tasks) introduces potential bias in quality of baseline documentation; beat-reference filtering for RE-Bench introduces fairness concerns about information parity.; Selection bias in paper choice (PaperBench subset of 23 ICML 2024 papers, RE-Bench subset of 7 tasks); Temporal bias (evaluation in 2026, rapidly evolving agent capabilities); Model-specific bias (evaluation primarily on Claude Sonnet 4.6; limited comparison with 4.5 base); Token budget constraints may advantage one format over another; Blinded evaluation mitigates some bias but judge models used for evaluation may have systematic preferences; Beat-reference filter on RE-Bench may introduce subtle bias in trace construction

Limitations

  • "Research whose contribution is a physical-world intervention (wet-lab biology, materials synthesis) falls outside this scope until the underlying experimental record is itself digitalised." The paper also notes that "on triton_cumsum and restricted_mlm the paper agent later overtakes via moves the trace does not name (an int8 kernel redesign and focused depth on a single architecture, respectively)," indicating that preserved failure traces can "constrain a capable agent from stepping outside the prior-run box depending on the agent's capabilities." The authors further note that "fabrication occurred in 2 baseline runs and 1 ARA run: structured artifacts reduce but do not eliminate hallucinated results."

Open questions raised

  • None explicitly stated as future work. The paper identifies limitations rather than gaps: the framework's scope excludes wet-lab biology and physical-world interventions; model-dependent performance suggests need for capability-aware artifact design; and the constraint effect of failure traces on weaker models warrants investigation.
  • Cross-artifact review via agent-to-agent ARA comparison (meaningful only as corpus reaches critical mass); scope limited to computation-expressible research; need for digitalization of physical experimental records in wet-lab biology and materials synthesis; integration with existing review ecosystems at scale; long-term sustainability of artifact infrastructure.
  • Existing efforts (FAIR principles, RO-Crate, Nanopublications, AGENTS.md) address fragments of the problem but "None of these efforts jointly structure scientific logic, executable code, and exploration history into a single operable object"
  • Cross-artifact review via agent-to-agent ARA comparison becomes meaningful only as corpus reaches critical mass
  • Physical-world interventions (wet-lab biology, materials synthesis) remain outside scope until experimental records are digitalized
  • Need for domain-specific ARA compilation pipelines beyond the generalized compiler approach
Data: PaperBench (23 ICML 2024 papers with 8,921 expert-annotated reproduction requirements); RE-Bench (5 of 7 tasks with METR MALT transcripts containing thousands of agent runs); METR eval-analysis-public dataset (24,008 agent runs across 21 frontier models); PaperBench (expert-annotated 8,921 reproduction requirements across 23 ICML 2024 papers); RE-Bench (5 tasks with METR MALT transcripts containing thousands of agent runs); METR eval-analysis-public dataset (24,008 agent runs across 21 frontier models). Code repository: github.com/AmberLJC/Agent-Native-Research-Artifact; PaperBench: 8,921 expert-annotated reproduction requirements across 23 ICML 2024 papers; RE-Bench: 5 open-ended extension tasks with METR MALT transcripts containing thousands of agent runs per task; METR eval-analysis-public dataset: 24,008 agent runs across 21 frontier modelsCode: github.com/AmberLJC/Agent-Native-Research-Artifact; https://github.com/AmberLJC/Agent-Native-Research-Artifact; github.com/AmberLJC/Agent-Native-Research-Artifact (main ARA implementation, Live Research Manager, ARA Compiler)Extracted from: pdfAgreement 45%

Explore related topics

Related papers