12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Vibe Researching as Wolf Coming: Can AI Agents with Skills Replace or Augment Social Scientists?

Yongjun Zhang · arXiv (Cornell University) · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
1/4
Quality (LMQS)
D
Evidence
0
Citations
0.00
FWCI

Methodology & findings

Study design

Conceptual framework development grounded in an operational system (scholar-skill).

Primary method

Operational system design and conceptual framework development grounded in artifact capabilities

Main result

The paper argues that "AI agents excel at speed, coverage, and methodological scaffolding but struggle with theoretical originality and tacit field knowledge." The study identifies a cognitive delegation boundary that "cuts through every stage of the research pipeline, not between stages," revealing that at each pipeline stage, some tasks are codifiable and delegable while others require tacit judgment. The author demonstrates through the scholar-skill system that "a single plugin can now cover 26 distinct research tasks from idea formalization to journal submission, coordinated by an orchestrator across 18 phases with 53 quality gates and five hard stops."

Research paradigm

Critical design inquiry with conceptual/theoretical framework development

Author conclusions

The paper concludes: "Vibe researching is already here; the profession's normative response is not." The author states that "The change in what a solo researcher can accomplish is real and significant" and emphasizes that "The delegation boundary is cognitive, not sequential. It cuts through every stage." Most urgently, the author argues: "Disclosure norms, pedagogical reform, equitable design, and deeper theory training are the four urgent interventions. We are not waiting for a future that may arrive. We are in it."

Risk of bias

Author-developed system creates potential for favorable bias toward the artifact; Single case study with no comparative analysis; Framework not empirically validated; System calibrated to 22 English-language journals, potentially introducing language and field bias; Author bias: case study system developed by the author; Scope limitation: framework not empirically validated; Generalizability: single tool focus may not represent broader AI research ecosystem; Single-system case study bias: Framework developed from author's own tool (scholar-skill), limiting generalizability; Lack of empirical validation: No user studies or controlled experiments to test framework predictions; Selection bias in journal calibration: System calibrated to 22 specific journals, potentially biasing recommendations toward those venues; Language bias: Training data and journal calibration skews English, disadvantaging non-English scholarship; Developer positionality: Author is the developer of scholar-skill, introducing potential confirmation bias; Taxonomy oversimplification: Four-type classification may not capture continuous nature of task characteristics

Limitations

  • The author explicitly states: "The case study is based on a single system (scholar-skill) developed by the author, which may not generalize to all AI research tools
  • The cognitive task framework, while grounded in the operational characteristics of the system, has not been empirically validated through user studies or controlled experiments
  • The four-type taxonomy simplifies a continuous space
  • some tasks may resist clean classification." Future work should "empirically test the framework's predictions about delegation effectiveness, conduct user studies comparing augmented and unaugmented research workflows, and examine how AI tool adoption varies across disciplines, institutions, and career stages."

Open questions raised

  • Empirical validation of the cognitive task framework through user studies or controlled experiments
  • Testing framework predictions about delegation effectiveness
  • Examination of how AI tool adoption varies across disciplines, institutions, and career stages
  • Study of AI's impact on academic labor, inequality, and knowledge production as social phenomena
  • Future work should: (1) empirically test the framework's predictions about delegation effectiveness; (2) conduct user studies comparing augmented and unaugmented research workflows; (3) examine how AI tool adoption varies across disciplines, institutions, and career stages; (4) study the impact of AI on academic stratification and labor market outcomes; (5) investigate pedagogical implications for graduate training.
  • Empirical validation of the cognitive task framework through user studies and controlled experiments
Extracted from: pdfAgreement 64%

Explore related topics

Related papers