12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Sound Agentic Science Requires Adversarial Experiments

Dionizije Fa, Marko Culjak · ArXiv.org · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
2/4
Quality (LMQS)
I
Evidence
0
Citations
0.00
FWCI

Methodology & findings

Study design

Position paper with illustrative experiment using two independent agents analyzing the same NHANES 2017-2018 dataset with opposite objectives.

Sample

2 groups

Primary method

Ordinary Least Squares (OLS) regression; specification variations including survey design adjustments, sample restrictions, covariate adjustment (age, sex, race/ethnicity), outcome definitions (PHQ-9 total vs. dichotomous at different thresholds), log transformations, quantile comparisons, and seasonal controls.

Main result

The study demonstrates that "Agent A finds a small but statistically significant negative association: for every 10 nmol/L higher serum 25(OH)D, PHQ-9 is 0.045 points lower on average (95% CI -0.068 to -0.023; p = 0.0006)" while "Agent B finds no evidence of an association: the estimated change is +0.0005 PHQ-9 points per 1 nmol/L (95% CI -0.0050 to +0.0060; p = 0.855)." This shows that opposing agents analyzing the same dataset with different prompts can produce conflicting yet defensible conclusions, illustrating how trivial it has become to generate plausible hypotheses.

Reports effect sizes and confidence intervals.

Research paradigm

Critical realism with emphasis on falsificationism (Popperian)

Author conclusions

"Without experimental evaluation and confirmation, agent outputs in empirical sciences should be treated as hypotheses rather than publishable conclusions. The more capable the agent, the more urgent this distinction becomes because capability increases the rate at which plausible analyses can be produced and selectively reinforced." The authors argue that "agentic science requires designing experiments that challenge the candidate hypotheses" rather than "further accelerating discovery" without engaging with reality.

Risk of bias

Publication bias favoring statistically significant results; Selective reporting and p-hacking in agent-assisted research; Confounding by indication in observational studies; Measurement error in biomedical data; Researcher degrees of freedom in analytic choices (specification choices, sample restrictions, covariate adjustment); Specification bias - multiple analytic choices leading to different conclusions; Publication bias - incentive structure optimizes for publishable results; p-hacking and selective reporting in traditional research practices; Selection bias from sample restrictions; Measurement error in biological data; Publication bias and p-hacking in scientific publishing (pre-existing condition); Selective reporting incentives in research; Confounding and measurement error in observational data; Arbitrary analytic choices masquerading as defensible specifications

Limitations

  • The paper acknowledges that "POPPER's current instantiation operates entirely on static databases
  • While its authors describe a general framework that could, in principle, incorporate physical experiments, no such deployment exists yet
  • Its experiments are statistical analyses of existing data, not physical interventions
  • The verification gap we identify, therefore, persists."

Open questions raised

  • Gap between verification capability in software engineering (rapid testing against specifications) versus empirical sciences (requires experimental validation)
  • Need for deployment of agentic frameworks like POPPER that incorporate physical experiments beyond static database analysis
  • Absence of established falsification-first standards in peer review for agent-assisted research
  • Lack of end-to-end agentic control over scientific workflows including automated laboratory experiments
  • Need for deployment of falsification-first frameworks (like POPPER) with access to physical experiments and automated laboratories
  • Gap between statistical falsification attempts and actual experimental validation
Data: NHANES 2017-2018 (National Health and Nutrition Examination Survey): XPT files (DEMO_J, DPQ_J, VID_J) merged by SEQN; NHANES 2017-2018; NHANES 2017-2018 (National Health and Nutrition Examination Survey, CDC & NCHS) - publicly available; data files: DEMO_J, DPQ_J, VID_JExtracted from: pdfAgreement 60%

Explore related topics

Related papers