AI scientists produce results without reasoning scientifically
Martiño Ríos-García, Nawaf Alampara, Chandan Gupta, Indrajeet Mandal, Sajid Mannan, Ali Asghar Aghajani et al. · ArXiv.org · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
Methodology & findings
Study design
Large-scale computational evaluation combining systematic performance analysis and behavioral analysis of epistemological structure.
Main result
The study found that "the base model is the primary determinant of both performance and behavior, accounting for 41.4% of explained variance versus 1.5% for the scaffold. Across all configurations, evidence is ignored in 68% of traces, refutation-driven belief revision occurs in 26%, and convergent multi-test evidence is rare." The authors further report that "agents routinely ignore evidence they have gathered, commit to hypotheses without testing them, and fail to revise beliefs when confronted with contradictory data."
Research paradigm
Empirical-computational evaluation of AI system behavior through systematic benchmarking and epistemological analysis
Author conclusions
"Current LLM-based agents execute scientific workflows but do not exhibit the epistemic patterns that characterize scientific reasoning. Outcomebased evaluation cannot detect these failures, and scaffold engineering alone cannot repair them. Until reasoning itself becomes a training target, the scientific knowledge produced by such agents cannot be justified by the process that generated it." The authors further conclude that "the tools scientists use shape the science they produce" and that "as long as they are evaluated only by the answers they produce, this difference will remain invisible, and it will shape the knowledge they help produce."
Risk of bias
Model selection bias: Only three frontier models evaluated (two commercial, one open-source); Annotation bias: Manual trace annotation by two experts using custom taxonomy; inter-rater reliability reported in Section H.5; Environmental bias: Domains selected may not represent full range of scientific inquiry; Temperature setting: All runs at temperature 0.0 may not reflect typical deployment conditions; Tool verbosity confound: Though controlled, tool description detail could influence performance; Selection bias in domain choice (eight specific domains may not represent all scientific reasoning types); Potential confounding: base model performance may reflect training data properties rather than reasoning capability; Annotation bias: manual annotation of reasoning traces by Claude Sonnet 4.5 (the model itself) could introduce circularity; Temperature setting of 0.0 may not reflect real-world deployment scenarios; Tool verbosity ablation may not capture all forms of scaffolding relevant to reasoning; Selection bias: Only three frontier models evaluated (two commercial, one open-source); results may not generalize to other LLM architectures; Temperature setting: All experiments used temperature 0.0, potentially not representative of real-world deployments with higher temperatures; Task design bias: Environments manually constructed; task difficulty and scope modulation may favor or disadvantage specific model capabilities; Annotation bias: Epistemological trace annotation performed by authors using custom taxonomy; potential for annotation drift despite inter-rater agreement checks; Verbosity confound: Tool verbosity controlled as ablation axis but may interact differently with different models; Iteration limits: Fixed iteration limits per environment may penalize models differently based on their convergence patterns
Limitations
- The authors note in their scope and limitations discussion (Section I) that "the framework has computational budget constraints affecting evaluation scale." They further state that "scaffold engineering alone cannot repair" the reasoning failures identified, and acknowledge that "outcome-based evaluation cannot detect these failures." The paper emphasizes that "the epistemic process by which they arrive at scientific conclusions is largely inaccessible to scrutiny" due to the statistical nature of LLM reasoning.
Open questions raised
- The authors identify several future research directions: (1) need for explicit measurement of epistemological behavior in AI systems, (2) requirement for training targets focused on reasoning process rather than outcome alone, (3) extension of evaluation methods beyond scientific agents to any domain where reasoning determines value, and (4) necessity of institutional structures enforcing epistemic norms for LLM-based agents analogous to peer review and replication in human science.
- Need for direct assessment of reasoning process in LLM-based agents rather than outcome-only evaluation
- Requirement for changes at the base-model training level to address epistemic reasoning patterns (scaffold engineering insufficient)
- Necessity for reasoning itself to become an explicit training target in LLM development
- Need for institutional structures enforcing epistemic norms in autonomous AI agents (analogous to peer review and replication in human science)
- Extension of evaluation framework beyond scientific agents to other domains where reasoning process determines result value
Explore related topics
Related papers
- Estimating the reproducibility of psychological scienceAlexander A. Aarts · 2015 · 8,669 citations
- Academic Integrity considerations of AI Large Language Models in the post-pandemic era: ChatGPT and beyondMike Perkins · 2023 · 668 citations
- Comparing scientific abstracts generated by ChatGPT to real abstracts with detectors and blinded human reviewersCatherine A. Gao · 2023 · 657 citations
- ChatGPT and the rise of generative AI: Threat to academic integrity?Damian Eke · 2023 · 476 citations
- Nonhuman “Authors” and Implications for the Integrity of Scientific Publication and Medical KnowledgeAnnette Flanagin · 2023 · 399 citations
- Fabrication and errors in the bibliographic citations generated by ChatGPTWilliam H. Walters · 2023 · 352 citations