The Last Human-Written Paper: Agent-Native Research Artifacts
Jiachen Liu, Jiaxin Pei, Jintao Huang, Chenglei Si, Ao Qu, Xiangru Tang et al. · ArXiv.org · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
Methodology & findings
Study design
Multi-layered mixed-method evaluation across three benchmarks (PaperBench, RE-Bench, METR eval-analysis-public dataset).
Primary method
Design Science with iterative refinement. Four-layer structured protocol design (Cognitive, Physical, Exploration Graph, Evidence layers) informed by taxonomy of reproduction-critical information categories from PaperBench rubrics.
Main result
ARA raises question-answering accuracy from 72.4% to 93.7% and reproduction success from 57.4% to 64.4%. On the understanding layer, "ARA outperforms the baseline at every category and every benchmark, with overall accuracy 93.7% vs. 72.4% (+21.3%) on 450 paired outcomes." On reproduction, "Across all 15 papers with complete paired runs (150 subtasks, 1,743 rubric requirements), ARA achieves a difficulty-weighted success rate of 64.4% vs. 57.4% for the baseline." On extension, "preserved failure traces in ARA accelerate progress, but can also constrain a capable agent from stepping outside the prior-run box depending on the agent's capabilities."
Research paradigm
Design science and artifact-centered research with agent-native computing
Author conclusions
The Agent-Native Research Artifact protocol recasts the primary research object from narrative document to agent-executable knowledge package. "Tolerable for human readers, these costs become critical when AI agents must understand, reproduce, and extend published work." The authors conclude that "research is scaling into a massively parallel enterprise where agents fork, extend, and merge each other's work at machine speed, shifting the bottleneck from individual productivity to artifact operability: narrative PDFs, compiled for sequential human reading, cannot be forked, diffed, or merged, but a structured, lossless artifact can, letting research compound like software."
Risk of bias
Selection bias in benchmark choice (ICML 2024 papers and RE-Bench tasks); potential evaluator bias in blinded judging; capability-dependent results (different models show inverted preferences); beat-reference filter applied to RE-Bench but fairness depends on filter correctness.; Selection bias in paper choice (only 15 of 23 PaperBench papers had companion repositories); task difficulty stratification may not be representative; baseline construction (LLM-synthesized paper writeups for RE-Bench tasks) introduces potential bias in quality of baseline documentation; beat-reference filtering for RE-Bench introduces fairness concerns about information parity.; Selection bias in paper choice (PaperBench subset of 23 ICML 2024 papers, RE-Bench subset of 7 tasks); Temporal bias (evaluation in 2026, rapidly evolving agent capabilities); Model-specific bias (evaluation primarily on Claude Sonnet 4.6; limited comparison with 4.5 base); Token budget constraints may advantage one format over another; Blinded evaluation mitigates some bias but judge models used for evaluation may have systematic preferences; Beat-reference filter on RE-Bench may introduce subtle bias in trace construction
Limitations
- "Research whose contribution is a physical-world intervention (wet-lab biology, materials synthesis) falls outside this scope until the underlying experimental record is itself digitalised." The paper also notes that "on triton_cumsum and restricted_mlm the paper agent later overtakes via moves the trace does not name (an int8 kernel redesign and focused depth on a single architecture, respectively)," indicating that preserved failure traces can "constrain a capable agent from stepping outside the prior-run box depending on the agent's capabilities." The authors further note that "fabrication occurred in 2 baseline runs and 1 ARA run: structured artifacts reduce but do not eliminate hallucinated results."
Open questions raised
- None explicitly stated as future work. The paper identifies limitations rather than gaps: the framework's scope excludes wet-lab biology and physical-world interventions; model-dependent performance suggests need for capability-aware artifact design; and the constraint effect of failure traces on weaker models warrants investigation.
- Cross-artifact review via agent-to-agent ARA comparison (meaningful only as corpus reaches critical mass); scope limited to computation-expressible research; need for digitalization of physical experimental records in wet-lab biology and materials synthesis; integration with existing review ecosystems at scale; long-term sustainability of artifact infrastructure.
- Existing efforts (FAIR principles, RO-Crate, Nanopublications, AGENTS.md) address fragments of the problem but "None of these efforts jointly structure scientific logic, executable code, and exploration history into a single operable object"
- Cross-artifact review via agent-to-agent ARA comparison becomes meaningful only as corpus reaches critical mass
- Physical-world interventions (wet-lab biology, materials synthesis) remain outside scope until experimental records are digitalized
- Need for domain-specific ARA compilation pipelines beyond the generalized compiler approach
Explore related topics
Related papers
- Preferred reporting items for systematic reviews and meta-analyses: the PRISMA statementDavid Moher · 2009 · 83,271 citations
- PRISMA 2020 explanation and elaboration: updated guidance and exemplars for reporting systematic reviewsMatthew J. Page · 2021 · 10,956 citations
- Systematic review of research on artificial intelligence applications in higher education – where are the educators?Olaf Zawacki‐Richter · 2019 · 5,282 citations
- State of the art and practice in AI in educationW. Holmes · 2022 · 758 citations
- Co-designing AI Education Curriculum with Cross-Disciplinary High School TeachersBenjamin Xie · 2024 · 28 citations
- GAIDeT (Generative AI Delegation Taxonomy): A taxonomy for humans to delegate tasks to generative artificial intelligence in scientific research and publishingYana Suchikova · 2025 · 24 citations