12,637 papers · updated 18 Sept 2026livingmeta.ai
← Browse all papers
AI evidence extraction

AI scientists produce results without reasoning scientifically

Martiño Ríos-García, Nawaf Alampara, Chandan Gupta, Indrajeet Mandal, Sajid Mannan, Ali Asghar Aghajani et al. · ArXiv.org · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
C
Evidence
0
Citations
0.00
FWCI

Methodology & findings

Study design

Large-scale computational evaluation combining systematic performance analysis and behavioral analysis of epistemological structure.

Main result

The study found that "the base model is the primary determinant of both performance and behavior, accounting for 41.4% of explained variance versus 1.5% for the scaffold. Across all configurations, evidence is ignored in 68% of traces, refutation-driven belief revision occurs in 26%, and convergent multi-test evidence is rare." The authors further report that "agents routinely ignore evidence they have gathered, commit to hypotheses without testing them, and fail to revise beliefs when confronted with contradictory data."

Research paradigm

Empirical-computational evaluation of AI system behavior through systematic benchmarking and epistemological analysis

Author conclusions

"Current LLM-based agents execute scientific workflows but do not exhibit the epistemic patterns that characterize scientific reasoning. Outcomebased evaluation cannot detect these failures, and scaffold engineering alone cannot repair them. Until reasoning itself becomes a training target, the scientific knowledge produced by such agents cannot be justified by the process that generated it." The authors further conclude that "the tools scientists use shape the science they produce" and that "as long as they are evaluated only by the answers they produce, this difference will remain invisible, and it will shape the knowledge they help produce."

Risk of bias

Model selection bias: Only three frontier models evaluated (two commercial, one open-source); Annotation bias: Manual trace annotation by two experts using custom taxonomy; inter-rater reliability reported in Section H.5; Environmental bias: Domains selected may not represent full range of scientific inquiry; Temperature setting: All runs at temperature 0.0 may not reflect typical deployment conditions; Tool verbosity confound: Though controlled, tool description detail could influence performance; Potential confounding: base model performance may reflect training data properties rather than reasoning capability; Annotation bias: manual annotation of reasoning traces by Claude Sonnet 4.5 (the model itself) could introduce circularity; Tool verbosity ablation may not capture all forms of scaffolding relevant to reasoning; results may not generalize to other LLM architectures; Temperature setting: All experiments used temperature 0.0, potentially not representative of real-world deployments with higher temperatures; Task design bias: Environments manually constructed; task difficulty and scope modulation may favor or disadvantage specific model capabilities; potential for annotation drift despite inter-rater agreement checks; Verbosity confound: Tool verbosity controlled as ablation axis but may interact differently with different models; Iteration limits: Fixed iteration limits per environment may penalize models differently based on their convergence patterns

Limitations

  • The authors note in their scope and limitations discussion (Section I) that "the framework has computational budget constraints affecting evaluation scale." They further state that "scaffold engineering alone cannot repair" the reasoning failures identified, and acknowledge that "outcome-based evaluation cannot detect these failures." The paper emphasizes that "the epistemic process by which they arrive at scientific conclusions is largely inaccessible to scrutiny" due to the statistical nature of LLM reasoning.

Open questions raised

  • The authors identify several future research directions:
  • need for explicit measurement of epistemological behavior in AI systems, (2) requirement for training targets focused on reasoning process rather than outcome alone, (3) extension of evaluation methods beyond scientific agents to any domain where reasoning determines value, and (4) necessity of institutional structures enforcing epistemic norms for LLM-based agents analogous to peer review and replication in human science.
  • Requirement for changes at the base-model training level to address epistemic reasoning patterns (scaffold engineering insufficient)
  • Gap between outcome-based evaluation (which cannot detect epistemic failures) and process-level evaluation
  • Investigation of how reasoning patterns adapt as new models and training approaches emerge
Data: https://huggingface.co/datasets/jablonkagroup/corral-environment-tasksCode: Interactive benchmark exploration: https://lamalab-org.github.io/corral/#environments; Corral framework (evaluation framework developed for this study)Extracted from: pdf

Explore related topics

Related papers