12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

AI for Auto-Research: Roadmap & User Guide

Lingdong Kong, Xian Sun, Wei Chow, Linfeng Li, Kevin Qinghong Lin, Xuan Billy Zhang et al. · ArXiv.org · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

10/10
Relevance
2/4
Quality (LMQS)
I
Evidence
0
Citations
0.00
FWCI

Methodology & findings

Study design

Systematic scoping review with multi-strategy literature collection: (1) Systematic keyword search across Google Scholar, Semantic Scholar, arXiv, and DBLP; (2) Snowball citation tracing (backward and forward); (3) Community and repository monitoring.

Primary method

Systematic literature review without quantitative meta-analysis. Qualitative synthesis using thematic organization across lifecycle phases. Benchmarking analysis and comparative assessment of system capabilities based on published benchmark results.

Main result

The study identifies "a sharp, stage-dependent boundary between reliable assistance and unreliable autonomy: AI excels at structured, retrieval-grounded, and tool-mediated tasks, but remains fragile for genuinely novel ideas, research-level experiments, and scientific judgment." Specifically, "AI capability is strongest when tasks are structured, grounded, and externally checkable, but drops sharply for open-ended research tasks requiring novelty, implicit domain knowledge, long-horizon reasoning, or scientific judgment." A critical finding shows that "Generated ideas often degrade after implementation, research code lags far behind pattern-matching benchmarks, and end-to-end autonomous systems have not yet consistently reached major-venue acceptance standards." The authors further observe that "80% of fully autonomous results are fabricated" in certain contexts.

Reports effect sizes.

Research paradigm

Systematic analysis of AI capabilities across research lifecycle stages; epistemological framework organizing four phases of academic research

Author conclusions

The authors conclude that "the most reliable deployment mode is human-governed collaboration rather than full autonomy: AI can reduce mechanical friction in retrieval, drafting, coding, visualization, review support, and dissemination, but researchers must retain responsibility for judgment, interpretation, experimental design, argumentation, and accountability." They further state that "effective systems increasingly rely on layered architectures that combine exploration, tool-based execution, and verification, suggesting that orchestration, provenance, and feedback design are as important as model scale." Finally, they argue that "AI use in research is becoming a governance problem rather than a detection problem: as AI assistance becomes routine, the key questions are disclosure, attribution, responsibility, and whether scientific integrity is preserved."

Risk of bias

Publication bias: commercial and proprietary dissemination tools underrepresented; open-source and published systems overrepresented; Domain bias: ML/NLP heavily overrepresented; limited coverage of chemistry, biology, physics workflows; Selection bias: inclusion restricted to papers with 'sufficient methodological or evaluative detail' may exclude early-stage work; Temporal bias: survey emphasizes 2023-2026 developments; foundational pre-2023 methods may be underweighted; Accessibility bias: closed systems excluded if 'insufficient technical or evaluative information is available'; Publication bias toward open-source and published systems; underrepresentation of commercial or closed systems; domain bias toward computer science and machine learning; maturity bias favoring well-benchmarked creation-stage tools over less-documented dissemination tools; Selection bias in literature collection toward computer science and machine learning domains; publication bias favoring benchmarked and open-sourced creation-stage tools over less-documented dissemination tools; potential gray literature gaps from commercial closed systems; geographic and language bias toward English-language publications on arXiv/Google Scholar.

Limitations

  • The authors state: "The resulting corpus spans all four phases of the lifecycle, but the distribution is uneven
  • Most documented systems concentrate on P1 (Creation), especially literature review, coding, and experiment automation, followed by P2 (Writing), P3 (Validation), and P4 (Dissemination)
  • This imbalance reflects both research maturity and publication availability: creation-stage tools are more frequently benchmarked and open-sourced, whereas dissemination-oriented tools are often commercial, workflow-specific, or evaluated through less standardized criteria." Additionally, the survey notes that "Nearly all benchmarks and systems target ML/NLP literature
  • cross-domain synthesis (chemistry, biology, physics) remains largely untested and likely requires domain-specific retrieval infrastructure."

Open questions raised

  • Faithfulness across phase boundaries: errors propagate downstream when AI systems generate plausible outputs without preserving evidence or provenance
  • Scientific judgment and novelty assessment: AI remains fragile for open-ended research requiring genuine novelty and implicit domain knowledge
  • Verification, reproducibility, and accountability: gap between artifact generation speed and verification speed
  • Citation, versioning, and source provenance: unclear attribution chains and version control for AI-assisted work
  • Governance, disclosure, and research integrity: lack of standards for AI use disclosure in academic publishing
  • Cross-domain generalization: systems concentrated on ML/NLP; limited infrastructure for chemistry, biology, physics
Data: Project Page: https://worldbench.github.io/awesome-ai-auto-research; GitHub Repo: https://github.com/worldbench/awesome-ai-auto-research; GitHub repository: https://github.com/worldbench/awesome-ai-auto-research; Multiple benchmarks listed in Table 2 including IdeaBench, LiveIdeaBench, AI Idea Bench 2025, ResearchBench, Scientist-Bench, LitSearch, DeepScholar-Bench, ReportBench, ScholarGym, SciNetBench, IDRBench, SWE-bench, MLAgentBench, LAB-Bench, DiscoveryBench, DiscoveryWorld, MLE-Bench, ScienceAgentBench, ResearchCodeBench, PaperBench, and many others; Project page maintained at: https://worldbench.github.io/awesome-ai-auto-researchCode: https://github.com/worldbench/awesome-ai-auto-research; SWE-agent (software engineering agent platform); OpenHands (open platform for software engineering agents); PaperCoder (paper-to-code translation); AIDE (ML engineering as tree search); Multiple tool implementations documented in Appendix A; GitHub Repo: https://github.com/worldbench/awesome-ai-auto-researchExtracted from: pdfAgreement 58%

Explore related topics

Related papers