12,637 papers · updated 18 Sept 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Plausible but Wrong: A case study on Agentic Failures in Astrophysical Workflows

Shivam Rawat, Lucie Flek · ArXiv.org · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
C
Evidence
0
Citations
0.00
FWCI

Methodology & findings

Study design

Structured empirical evaluation case study using controlled computational tasks with quantitative metrics (ESR, PAS, NAS) for tool-grounded computation and qualitative assessment (parameter recovery score, physical plausibility) for multi-step inference tasks.

Main result

The study found that "in the One-Shot setting, access to domain-specific context yields an approximately ∼6× performance improvement (0.85 vs. ≈0 without context), with the primary failure mode being silent incorrect computation—syntactically valid code that produces plausible but inaccurate results." In the Deep Research setting, "the system frequently exhibits silent failures across stress tests, producing physically inconsistent posteriors without self-diagnosis."

Research paradigm

Empirical evaluation of computational system reliability under realistic conditions

Author conclusions

The authors conclude that "agentic systems do not primarily fail by crashing—they fail by producing confident, incorrect results or by silently breaking pipelines without diagnosis." They emphasize that "the most concerning failure mode in agentic scientific workflows is not overt failure, but confident generation of incorrect results." Furthermore, they state: "we argue that rigorous reliability evaluation is a prerequisite for deploying agentic AI in scientific workflows, where undetected errors can directly compromise scientific conclusions."

Risk of bias

Selection bias: Only astrophysical tasks used; findings may not generalize to other scientific domains; System-specific bias: Evaluation limited to CMBAgent; other agentic systems may exhibit different failure modes; Model selection bias: CMBAgent configured with GPT-4o-mini baseline; different LLMs may perform differently; Task design bias: Research-driven tasks explicitly designed as stress tests; may not reflect typical usage patterns; Reference bias: Ground truth for Deep Research tasks drawn from published literature; observational biases may persist; Single system evaluated (CMBAgent) - results may not generalize to other agentic systems; Limited sample size for Deep Research tasks (N=5 trials per task); Evaluator bias in qualitative assessment of physical plausibility; Reference solutions may contain domain knowledge that newer models lack; Domain specificity: Tasks limited to astrophysics; Evaluation metric design: Weighted parameter scoring (PAS) may introduce bias if weights do not align with true scientific importance

Limitations

  • The evaluation is limited to a single agentic system (CMBAgent) applied to astrophysical inference tasks
  • "Evaluation of this workflow is therefore qualitative, based on parameter recovery and physical consistency of the reported outputs rather than statistical aggregation across trials" for the Deep Research workflow
  • The study focuses exclusively on two fully automated workflows and "excludes the Human in the Loop mode as user feedback introduces variability that complicates reproducible benchmarking." The evaluation does not propose or critique agentic architectures themselves but rather evaluates reliability when deployed in realistic workflows.

Open questions raised

  • Systematic analysis of reliability, failure modes, and error propagation in multi-step scientific pipelines remains largely unexplored
  • Limited understanding of agentic system behavior under realistic scientific workflow conditions compared to synthetic benchmarks
  • Need for structured evaluation frameworks that distinguish between execution success, parameter accuracy, and numerical fidelity
  • Gap in characterizing how silent failures propagate through complex multi-stage workflows with external tool interaction
  • Limited work on failure mode transparency—whether systems can self-diagnose and report known pathologies
  • Distinction between visible and silent failures in agentic systems needs characterization
Data: Union2.1 Type Ia supernova distance-redshift data (referenced: Suzuki et al., 2012); NGC 3198 galaxy rotation curve data (SPARC; referenced: Karukes et al., 2015); NASA Exoplanet Archive data filtered for planets with Mp > 2 M⊕; SLACS strong lensing systems data (referenced: Bolton et al., 2008); CMBAgent benchmark repository (https://github.com/cmbagent/Benchmarks); NASA Exoplanet Archive data (referenced as /home/sr/Desktop/code/cmbagent/cmbagent_systematics/deepresearch/task/NASA_exoplanet_archive.csv)Code: CMBAgent repository: https://github.com/cmbagent/Benchmarks (Contributors, 2024); CMBAgent source code: arXiv:2507.07257 (Xu et al., 2025); Evaluation framework released by authors (mentioned in abstract as 'We release our evaluation framework')Extracted from: pdf

Explore related topics

Related papers