Plausible but Wrong: A case study on Agentic Failures in Astrophysical Workflows
Shivam Rawat, Lucie Flek · ArXiv.org · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
Methodology & findings
Study design
Structured empirical evaluation case study using controlled computational tasks with quantitative metrics (ESR, PAS, NAS) for tool-grounded computation and qualitative assessment (parameter recovery score, physical plausibility) for multi-step inference tasks.
Main result
The study found that "in the One-Shot setting, access to domain-specific context yields an approximately ∼6× performance improvement (0.85 vs. ≈0 without context), with the primary failure mode being silent incorrect computation—syntactically valid code that produces plausible but inaccurate results." In the Deep Research setting, "the system frequently exhibits silent failures across stress tests, producing physically inconsistent posteriors without self-diagnosis."
Research paradigm
Empirical evaluation of computational system reliability under realistic conditions
Author conclusions
The authors conclude that "agentic systems do not primarily fail by crashing—they fail by producing confident, incorrect results or by silently breaking pipelines without diagnosis." They emphasize that "the most concerning failure mode in agentic scientific workflows is not overt failure, but confident generation of incorrect results." Furthermore, they state: "we argue that rigorous reliability evaluation is a prerequisite for deploying agentic AI in scientific workflows, where undetected errors can directly compromise scientific conclusions."
Risk of bias
Selection bias: Only astrophysical tasks used; findings may not generalize to other scientific domains; System-specific bias: Evaluation limited to CMBAgent; other agentic systems may exhibit different failure modes; Model selection bias: CMBAgent configured with GPT-4o-mini baseline; different LLMs may perform differently; Task design bias: Research-driven tasks explicitly designed as stress tests; may not reflect typical usage patterns; Reference bias: Ground truth for Deep Research tasks drawn from published literature; observational biases may persist; Single system evaluated (CMBAgent) - results may not generalize to other agentic systems; Limited sample size for Deep Research tasks (N=5 trials per task); Task design bias - tasks selected from CMBAgent benchmark repository; Evaluator bias in qualitative assessment of physical plausibility; Reference solutions may contain domain knowledge that newer models lack; Single system evaluation: Only CMBAgent is evaluated; results may not generalize to other agentic systems; Selection bias in task design: Tasks are selected from CMBAgent benchmark repository, potentially favoring the system; Limited trial size for Deep Research: Only 5 trials per task may be insufficient for robust statistical characterization; Domain specificity: Tasks limited to astrophysics; generalizability to other scientific domains unclear; Evaluation metric design: Weighted parameter scoring (PAS) may introduce bias if weights do not align with true scientific importance
Limitations
- The evaluation is limited to a single agentic system (CMBAgent) applied to astrophysical inference tasks
- "Evaluation of this workflow is therefore qualitative, based on parameter recovery and physical consistency of the reported outputs rather than statistical aggregation across trials" for the Deep Research workflow
- The study focuses exclusively on two fully automated workflows and "excludes the Human in the Loop mode as user feedback introduces variability that complicates reproducible benchmarking." The evaluation does not propose or critique agentic architectures themselves but rather evaluates reliability when deployed in realistic workflows.
Open questions raised
- Systematic analysis of reliability, failure modes, and error propagation in multi-step scientific pipelines remains largely unexplored
- Limited understanding of agentic system behavior under realistic scientific workflow conditions compared to synthetic benchmarks
- Need for structured evaluation frameworks that distinguish between execution success, parameter accuracy, and numerical fidelity
- Gap in characterizing how silent failures propagate through complex multi-stage workflows with external tool interaction
- Limited work on failure mode transparency—whether systems can self-diagnose and report known pathologies
- Systematic reliability evaluation in realistic multi-step scientific workflows remains underexplored
Explore related topics
Related papers
- Estimating the reproducibility of psychological scienceAlexander A. Aarts · 2015 · 8,669 citations
- Academic Integrity considerations of AI Large Language Models in the post-pandemic era: ChatGPT and beyondMike Perkins · 2023 · 668 citations
- Comparing scientific abstracts generated by ChatGPT to real abstracts with detectors and blinded human reviewersCatherine A. Gao · 2023 · 657 citations
- ChatGPT and the rise of generative AI: Threat to academic integrity?Damian Eke · 2023 · 476 citations
- Nonhuman “Authors” and Implications for the Integrity of Scientific Publication and Medical KnowledgeAnnette Flanagin · 2023 · 399 citations
- Fabrication and errors in the bibliographic citations generated by ChatGPTWilliam H. Walters · 2023 · 352 citations