12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Many AI Analysts, One Dataset: Navigating the Agentic Data Science Multiverse

Martín Bertrán, Riccardo Fogliato, Zhiwei Steven Wu · arXiv (Cornell University) · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
3/4
Quality (LMQS)
E
Evidence
0
Citations
0.00
FWCI

Methodology & findings

Study design

Fully autonomous AI analyst agents (ReAct tool-using agents) deployed across 3 datasets and 4 LLMs with 5 analyst personas (standard, negative, positive, confirmation-seeking, strong confirmation-seeking), totaling approximately 5,000 runs.

Sample

N = 5000, 6 groups

Primary method

Ordinary least squares (OLS) regression, logistic regression, linear probability models, generalized additive models (GAMs), mixed-effects models, tree/ensemble methods, regularized regression, Bayesian variants. Variance estimators: heteroskedasticity-consistent (HC) variants, robust clustering, bootstrap methods. Inference: two-sided p-values at α=0.05 for standard analysis; one-sided tests under confirmation-seeking personas. Software: Python with persistent sessions; specific packages not explicitly named but analysis code generated by AI agents.

Main result

The study found that "across all three datasets, AI analysts display wide dispersion in effect sizes, p-values, and binary support decisions; independent runs frequently reverse whether the hypothesis is judged supported." Furthermore, "the dispersion is steerable: changing the analyst persona or LLM systematically shifts the outcome distribution even among judge-approved analyses. Comparing the most skeptical persona to the most confirmation-seeking yields support-rate differences of 34 to 66 percentage points across datasets."

Reports effect sizes and confidence intervals.

Research paradigm

Empirical-computational (mixed-methods: automated experimentation with human-centered metascience)

Author conclusions

"The central challenge is not that automated analyses are wrong but that they are abundant. For a fixed dataset and hypothesis, AI analysts produce many defensible pipelines that reach meaningfully different conclusions: a structural vulnerability to selective reporting." The authors advocate that "specification curves and multiverse reporting should accompany any AI-generated analysis, and we argue that the exact prompts used should be disclosed as part of the methodological specification, on par with code and data." They also note that "the same agentic scale that creates vulnerability can make multiverse mapping routine," enabling "computational stress tests" for published findings and making "analytic variability from a hidden liability into a visible, measurable quantity."

Risk of bias

LLM hallucination and confabulation (documented in pilot runs with fully hallucinated results); Training data contamination: soccer dataset widely known in training corpora, creating high-contamination benchmark; Persona-induced steering of analytical conclusions; Shared blind spots in AI models inherited from training data; Prompt sensitivity and sycophancy in LLMs; Auditor bias: LLM-as-judge methods carry known biases (acknowledged in related work); Specification search bias under confirmation-seeking conditions; LLM training data contamination (high contamination risk for soccer dataset due to wide publication of Silberzahn et al. 2018 study); AI auditor bias and subjectivity in defining 'reasonable' analyses; LLM shared blind spots and modeling convention biases inherited from training data; Persona-driven steering of analytical conclusions through prompt manipulation; Hallucination risk in LLM-generated analyses (noted in pilot runs); Directionality confusion in fewer than 1% of analyses (corrected via post-processing); Data contamination risk: soccer dataset and conclusions widely known (high training data contamination); metr-rct recent and less likely in training corpora; anes-views low contamination; LLM hallucination and false reporting: pilot runs showed some analyses produced 'confident reports with fully hallucinated results'; Training data recall: some analyses 'recalled published findings' rather than analyzing provided dataset; Shared blind spots: LLMs may inherit systematic biases from training data; Prompt sensitivity: different framing shifts conclusions substantially; Model-specific biases: different LLMs produce different conclusion distributions

Limitations

  • The authors state: "Even with a pre-specified estimand, the space of defensible analyses remains large enough that comprehensive human review is impractical
  • Automated auditing is therefore necessary, but any definition of a 'reasonable' analysis depends on chosen standards, and how best to evaluate LLM-based auditors remains an open question." Additionally, "the analytical paths AI analysts explore need not coincide with those human analysts would traverse: LLMs may favor different modeling conventions, overlook domain-specific considerations, or exhibit shared blind spots inherited from training data
  • The multiverse we observe is therefore an AI-generated one, and its overlap with the human multiverse remains an empirical question."

Open questions raised

  • The overlap between the AI-generated multiverse and the human multiverse remains an empirical question
  • How best to evaluate LLM-based auditors as validity checkers
  • The broader implications of AI-assisted analysis beyond fully autonomous agents, including human-AI collaborative workflows
  • Methodological standards for defining 'reasonable' specifications in the context of AI analysis
  • Empirical validation of overlap between AI-generated multiverse and human analyst multiverse
  • Best practices for evaluating LLM-based auditors and defining 'reasonable' analyses
Data: soccer: Soccer referee bias dataset from Silberzahn et al. [2018] (publicly available); metr-rct: METR coding RCT dataset from Becker et al. [2025] (recent, described as unlikely to appear in training corpora); anes-views: American National Election Studies Time Series Cumulative File (1948-2020), [American National Election Studies, 2025]; Soccer referee bias dataset (Silberzahn et al., 2018); METR-RCT: AI-assisted programming randomized controlled trial (Becker et al., 2025); ANES (American National Election Studies) Time Series Cumulative Data File (1948-2020); soccer (soccer referees dataset from Silberzahn et al., 2018); metr-rct (METR coding RCT from Becker et al., 2025); anes-views (American National Election Studies Time Series Cumulative File, 1948-2020)Code: Not explicitly mentioned; reproducible code was produced by each AI analyst but no central repository link provided; Inspect AI framework (AI Security Institute) - used to implement ReAct agentsExtracted from: pdfAgreement 48%

Explore related topics

Related papers