12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Plausible but Wrong: A case study on Agentic Failures in Astrophysical Workflows

Shivam Rawat, Lucie Flek · ArXiv.org · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
2/4
Quality (LMQS)
C
Evidence
0
Citations
0.00
FWCI

Methodology & findings

Study design

Structured empirical evaluation case study using controlled computational tasks with quantitative metrics (ESR, PAS, NAS) for tool-grounded computation and qualitative assessment (parameter recovery score, physical plausibility) for multi-step inference tasks.

Main result

The study found that "in the One-Shot setting, access to domain-specific context yields an approximately ∼6× performance improvement (0.85 vs. ≈0 without context), with the primary failure mode being silent incorrect computation—syntactically valid code that produces plausible but inaccurate results." In the Deep Research setting, "the system frequently exhibits silent failures across stress tests, producing physically inconsistent posteriors without self-diagnosis."

Research paradigm

Empirical evaluation of computational system reliability under realistic conditions

Author conclusions

The authors conclude that "agentic systems do not primarily fail by crashing—they fail by producing confident, incorrect results or by silently breaking pipelines without diagnosis." They emphasize that "the most concerning failure mode in agentic scientific workflows is not overt failure, but confident generation of incorrect results." Furthermore, they state: "we argue that rigorous reliability evaluation is a prerequisite for deploying agentic AI in scientific workflows, where undetected errors can directly compromise scientific conclusions."

Risk of bias

Selection bias: Only astrophysical tasks used; findings may not generalize to other scientific domains; System-specific bias: Evaluation limited to CMBAgent; other agentic systems may exhibit different failure modes; Model selection bias: CMBAgent configured with GPT-4o-mini baseline; different LLMs may perform differently; Task design bias: Research-driven tasks explicitly designed as stress tests; may not reflect typical usage patterns; Reference bias: Ground truth for Deep Research tasks drawn from published literature; observational biases may persist; Single system evaluated (CMBAgent) - results may not generalize to other agentic systems; Limited sample size for Deep Research tasks (N=5 trials per task); Task design bias - tasks selected from CMBAgent benchmark repository; Evaluator bias in qualitative assessment of physical plausibility; Reference solutions may contain domain knowledge that newer models lack; Single system evaluation: Only CMBAgent is evaluated; results may not generalize to other agentic systems; Selection bias in task design: Tasks are selected from CMBAgent benchmark repository, potentially favoring the system; Limited trial size for Deep Research: Only 5 trials per task may be insufficient for robust statistical characterization; Domain specificity: Tasks limited to astrophysics; generalizability to other scientific domains unclear; Evaluation metric design: Weighted parameter scoring (PAS) may introduce bias if weights do not align with true scientific importance

Limitations

  • The evaluation is limited to a single agentic system (CMBAgent) applied to astrophysical inference tasks
  • "Evaluation of this workflow is therefore qualitative, based on parameter recovery and physical consistency of the reported outputs rather than statistical aggregation across trials" for the Deep Research workflow
  • The study focuses exclusively on two fully automated workflows and "excludes the Human in the Loop mode as user feedback introduces variability that complicates reproducible benchmarking." The evaluation does not propose or critique agentic architectures themselves but rather evaluates reliability when deployed in realistic workflows.

Open questions raised

  • Systematic analysis of reliability, failure modes, and error propagation in multi-step scientific pipelines remains largely unexplored
  • Limited understanding of agentic system behavior under realistic scientific workflow conditions compared to synthetic benchmarks
  • Need for structured evaluation frameworks that distinguish between execution success, parameter accuracy, and numerical fidelity
  • Gap in characterizing how silent failures propagate through complex multi-stage workflows with external tool interaction
  • Limited work on failure mode transparency—whether systems can self-diagnose and report known pathologies
  • Systematic reliability evaluation in realistic multi-step scientific workflows remains underexplored
Data: Union2.1 Type Ia supernova distance-redshift data (referenced: Suzuki et al., 2012); NGC 3198 galaxy rotation curve data (SPARC; referenced: Karukes et al., 2015); NASA Exoplanet Archive data filtered for planets with Mp > 2 M⊕; SLACS strong lensing systems data (referenced: Bolton et al., 2008); CMBAgent benchmark repository (https://github.com/cmbagent/Benchmarks); Union2.1 Type Ia supernova distance-redshift data; NGC 3198 galaxy rotation curve data (SPARC); NASA Exoplanet Archive data; SLACS strong-lensing survey data; Union2.1 Type Ia supernova distance–redshift data (referenced as /home/sr/Desktop/code/cmbagent/cmbagent_systematics/deepresearch/task/SCPUnion2.1_mu_vs_z.txt); NGC 3198 galaxy rotation curve data (SPARC archive, referenced as /home/sr/Desktop/code/cmbagent/cmbagent_systematics/deepresearch/task/NGC3198_rotmod.txt); NASA Exoplanet Archive data (referenced as /home/sr/Desktop/code/cmbagent/cmbagent_systematics/deepresearch/task/NASA_exoplanet_archive.csv); SLACS strong-lensing survey data (referenced as /home/sr/Desktop/code/cmbagent/cmbagent_systematics/deepresearch/task/Slac_data.csv); CMBAgent benchmark repository datasets (Contributors, 2024; https://github.com/cmbagent/Benchmarks)Code: CMBAgent repository: https://github.com/cmbagent/Benchmarks (Contributors, 2024); CMBAgent source code: arXiv:2507.07257 (Xu et al., 2025); CMBAgent repository (Xu et al., 2025) - https://github.com/cmbagent/Benchmarks; Evaluation framework released by authors (mentioned in abstract as 'We release our evaluation framework'); CMBAgent repository: https://github.com/cmbagent/Benchmarks (CMBAgent benchmark repository, Contributors 2024); Full CMBAgent framework: arXiv:2507.07257 (Xu et al., 2025) - details on agent-to-model assignments provided in repositoryExtracted from: pdfAgreement 39%

Explore related topics

Related papers