12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

AI scientists produce results without reasoning scientifically

Martiño Ríos-García, Nawaf Alampara, Chandan Gupta, Indrajeet Mandal, Sajid Mannan, Ali Asghar Aghajani et al. · ArXiv.org · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
2/4
Quality (LMQS)
C
Evidence
0
Citations
0.00
FWCI

Methodology & findings

Study design

Large-scale computational evaluation combining systematic performance analysis and behavioral analysis of epistemological structure.

Main result

The study found that "the base model is the primary determinant of both performance and behavior, accounting for 41.4% of explained variance versus 1.5% for the scaffold. Across all configurations, evidence is ignored in 68% of traces, refutation-driven belief revision occurs in 26%, and convergent multi-test evidence is rare." The authors further report that "agents routinely ignore evidence they have gathered, commit to hypotheses without testing them, and fail to revise beliefs when confronted with contradictory data."

Research paradigm

Empirical-computational evaluation of AI system behavior through systematic benchmarking and epistemological analysis

Author conclusions

"Current LLM-based agents execute scientific workflows but do not exhibit the epistemic patterns that characterize scientific reasoning. Outcomebased evaluation cannot detect these failures, and scaffold engineering alone cannot repair them. Until reasoning itself becomes a training target, the scientific knowledge produced by such agents cannot be justified by the process that generated it." The authors further conclude that "the tools scientists use shape the science they produce" and that "as long as they are evaluated only by the answers they produce, this difference will remain invisible, and it will shape the knowledge they help produce."

Risk of bias

Model selection bias: Only three frontier models evaluated (two commercial, one open-source); Annotation bias: Manual trace annotation by two experts using custom taxonomy; inter-rater reliability reported in Section H.5; Environmental bias: Domains selected may not represent full range of scientific inquiry; Temperature setting: All runs at temperature 0.0 may not reflect typical deployment conditions; Tool verbosity confound: Though controlled, tool description detail could influence performance; Selection bias in domain choice (eight specific domains may not represent all scientific reasoning types); Potential confounding: base model performance may reflect training data properties rather than reasoning capability; Annotation bias: manual annotation of reasoning traces by Claude Sonnet 4.5 (the model itself) could introduce circularity; Temperature setting of 0.0 may not reflect real-world deployment scenarios; Tool verbosity ablation may not capture all forms of scaffolding relevant to reasoning; Selection bias: Only three frontier models evaluated (two commercial, one open-source); results may not generalize to other LLM architectures; Temperature setting: All experiments used temperature 0.0, potentially not representative of real-world deployments with higher temperatures; Task design bias: Environments manually constructed; task difficulty and scope modulation may favor or disadvantage specific model capabilities; Annotation bias: Epistemological trace annotation performed by authors using custom taxonomy; potential for annotation drift despite inter-rater agreement checks; Verbosity confound: Tool verbosity controlled as ablation axis but may interact differently with different models; Iteration limits: Fixed iteration limits per environment may penalize models differently based on their convergence patterns

Limitations

  • The authors note in their scope and limitations discussion (Section I) that "the framework has computational budget constraints affecting evaluation scale." They further state that "scaffold engineering alone cannot repair" the reasoning failures identified, and acknowledge that "outcome-based evaluation cannot detect these failures." The paper emphasizes that "the epistemic process by which they arrive at scientific conclusions is largely inaccessible to scrutiny" due to the statistical nature of LLM reasoning.

Open questions raised

  • The authors identify several future research directions: (1) need for explicit measurement of epistemological behavior in AI systems, (2) requirement for training targets focused on reasoning process rather than outcome alone, (3) extension of evaluation methods beyond scientific agents to any domain where reasoning determines value, and (4) necessity of institutional structures enforcing epistemic norms for LLM-based agents analogous to peer review and replication in human science.
  • Need for direct assessment of reasoning process in LLM-based agents rather than outcome-only evaluation
  • Requirement for changes at the base-model training level to address epistemic reasoning patterns (scaffold engineering insufficient)
  • Necessity for reasoning itself to become an explicit training target in LLM development
  • Need for institutional structures enforcing epistemic norms in autonomous AI agents (analogous to peer review and replication in human science)
  • Extension of evaluation framework beyond scientific agents to other domains where reasoning process determines result value
Data: https://huggingface.co/datasets/jablonkagroup/corral-environment-tasks; https://huggingface.co/datasets/jablonkagroup/corral_runs_reports; https://huggingface.co/datasets/jablonkagroup/corral-traces; https://huggingface.co/datasets/jablonkagroup/corral-QAs; https://huggingface.co/datasets/jablonkagroup/corral-QAs-topic_reports; https://huggingface.co/datasets/jablonkagroup/corral-QAs-reports; https://huggingface.co/datasets/jablonkagroup/corral-oss-trace-logprobs; https://huggingface.co/datasets/jablonkagroup/rise_ai_scientists; https://huggingface.co/datasets/jablonkagroup/questions4manual_annotation; https://huggingface.co/datasets/jablonkagroup/corral-reasoning-annotations; https://huggingface.co/datasets/jablonkagroup/corral_lfm_binomial_results; https://huggingface.co/datasets/jablonkagroup/corral-intervention-traces; https://huggingface.co/datasets/jablonkagroup/corral-intervention-reports; corral-environment-tasks: https://huggingface.co/datasets/jablonkagroup/corral-environment-tasks; corral_runs_reports: https://huggingface.co/datasets/jablonkagroup/corral_runs_reports; corral-traces: https://huggingface.co/datasets/jablonkagroup/corral-traces; corral-QAs: https://huggingface.co/datasets/jablonkagroup/corral-QAs; corral-QAs-topic_reports: https://huggingface.co/datasets/jablonkagroup/corral-QAs-topic_reports; corral-QAs-reports: https://huggingface.co/datasets/jablonkagroup/corral-QAs-reports; corral-oss-trace-logprobs: https://huggingface.co/datasets/jablonkagroup/corral-oss-trace-logprobs; rise_ai_scientists: https://huggingface.co/datasets/jablonkagroup/rise_ai_scientists; questions4manual_annotation: https://huggingface.co/datasets/jablonkagroup/questions4manual_annotation; corral-reasoning-annotations: https://huggingface.co/datasets/jablonkagroup/corral-reasoning-annotations; corral_lfm_binomial_results: https://huggingface.co/datasets/jablonkagroup/corral_lfm_binomial_results; corral-intervention-traces: https://huggingface.co/datasets/jablonkagroup/corral-intervention-traces; corral-intervention-reports: https://huggingface.co/datasets/jablonkagroup/corral-intervention-reportsCode: Interactive benchmark exploration: https://lamalab-org.github.io/corral/#environments; Annotated trace browser: https://lamalab-org.github.io/corral/#explainers; Interactive environments and explainers browsable at https://lamalab-org.github.io/corral/#environments; Annotated traces browsable at https://lamalab-org.github.io/corral/#explainers; Corral framework (evaluation framework developed for this study); https://lamalab-org.github.io/corral/#environments (interactive exploration of environments); https://lamalab-org.github.io/corral/#explainers (browsable annotated traces)Extracted from: pdfAgreement 38%

Explore related topics

Related papers