Sound Agentic Science Requires Adversarial Experiments
Dionizije Fa, Marko Culjak · ArXiv.org · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
Methodology & findings
Study design
Position paper with illustrative experiment using two independent agents analyzing the same NHANES 2017-2018 dataset with opposite objectives.
Sample
2 groups
Primary method
Ordinary Least Squares (OLS) regression; specification variations including survey design adjustments, sample restrictions, covariate adjustment (age, sex, race/ethnicity), outcome definitions (PHQ-9 total vs. dichotomous at different thresholds), log transformations, quantile comparisons, and seasonal controls.
Main result
The study demonstrates that "Agent A finds a small but statistically significant negative association: for every 10 nmol/L higher serum 25(OH)D, PHQ-9 is 0.045 points lower on average (95% CI -0.068 to -0.023; p = 0.0006)" while "Agent B finds no evidence of an association: the estimated change is +0.0005 PHQ-9 points per 1 nmol/L (95% CI -0.0050 to +0.0060; p = 0.855)." This shows that opposing agents analyzing the same dataset with different prompts can produce conflicting yet defensible conclusions, illustrating how trivial it has become to generate plausible hypotheses.
Reports effect sizes and confidence intervals.
Research paradigm
Critical realism with emphasis on falsificationism (Popperian)
Author conclusions
"Without experimental evaluation and confirmation, agent outputs in empirical sciences should be treated as hypotheses rather than publishable conclusions. The more capable the agent, the more urgent this distinction becomes because capability increases the rate at which plausible analyses can be produced and selectively reinforced." The authors argue that "agentic science requires designing experiments that challenge the candidate hypotheses" rather than "further accelerating discovery" without engaging with reality.
Risk of bias
Publication bias favoring statistically significant results; Selective reporting and p-hacking in agent-assisted research; Confounding by indication in observational studies; Measurement error in biomedical data; Researcher degrees of freedom in analytic choices (specification choices, sample restrictions, covariate adjustment); Specification bias - multiple analytic choices leading to different conclusions; Publication bias - incentive structure optimizes for publishable results; p-hacking and selective reporting in traditional research practices; Selection bias from sample restrictions; Measurement error in biological data; Publication bias and p-hacking in scientific publishing (pre-existing condition); Selective reporting incentives in research; Confounding and measurement error in observational data; Arbitrary analytic choices masquerading as defensible specifications
Limitations
- The paper acknowledges that "POPPER's current instantiation operates entirely on static databases
- While its authors describe a general framework that could, in principle, incorporate physical experiments, no such deployment exists yet
- Its experiments are statistical analyses of existing data, not physical interventions
- The verification gap we identify, therefore, persists."
Open questions raised
- Gap between verification capability in software engineering (rapid testing against specifications) versus empirical sciences (requires experimental validation)
- Need for deployment of agentic frameworks like POPPER that incorporate physical experiments beyond static database analysis
- Absence of established falsification-first standards in peer review for agent-assisted research
- Lack of end-to-end agentic control over scientific workflows including automated laboratory experiments
- Need for deployment of falsification-first frameworks (like POPPER) with access to physical experiments and automated laboratories
- Gap between statistical falsification attempts and actual experimental validation
Explore related topics
Related papers
- Estimating the reproducibility of psychological scienceAlexander A. Aarts · 2015 · 8,669 citations
- Academic Integrity considerations of AI Large Language Models in the post-pandemic era: ChatGPT and beyondMike Perkins · 2023 · 668 citations
- Comparing scientific abstracts generated by ChatGPT to real abstracts with detectors and blinded human reviewersCatherine A. Gao · 2023 · 657 citations
- ChatGPT in education: Strategies for responsible implementationMohanad Halaweh · 2023 · 576 citations
- ChatGPT and the rise of generative AI: Threat to academic integrity?Damian Eke · 2023 · 476 citations
- Nonhuman “Authors” and Implications for the Integrity of Scientific Publication and Medical KnowledgeAnnette Flanagin · 2023 · 399 citations