From Specification to Execution: AI Assisted Scientific Workflow Management
Komal Thareja, Hamza Safri, Rajiv Mayani, Anirban Mandal, Ewa Deelman · ArXiv.org · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
Methodology & findings
Study design
Case study with comparative evaluation.
Primary method
Design science with specification-driven development; system architecture design integrating multiple components (AI authoring, debugging agent, Pegasus WMS, MCP interface)
Main result
The system generated and executed large-scale workflows with thousands of jobs, reduced debugging effort, and allowed non-expert users to construct workflows with expert-level design patterns. "The skill-based approach brings non-expert users closer to expert-level workflow design, reducing the need for manual expertise but leaving room for expert refinement." Claude produced a production-ready workflow in 2 sessions (52 prompts, ~$10–15) with superior reproducibility compared to Codex and Kimi, which required multiple sessions and exhibited convergence issues with Pegasus-specific staging conflicts.
Research paradigm
Design science / pragmatist
Author conclusions
"We presented an AI-assisted approach to scientific workflow management that combines specification-driven workflow generation, automated debugging, hierarchical execution with Pegasus WMS, and MCP-based remote management." The authors conclude that "the system could generate and autonomously execute workflows with thousands of jobs" and that "workflow structure, in particular the number of rounds, is the dominant factor in execution cost." They propose future work on "evaluating the performance of the MCP layer, and building an AI-driven platform that guides users in designing more complex workflows, and manages the full workflow lifecycle."
Risk of bias
Selection bias in use case choice (federated learning may not represent all workflow types); Comparison bias: manual baseline developed by single expert may not represent typical manual workflow development; Funding bias: authors acknowledge NSF support but work is on their own system; LLM selection bias: Claude, Codex, and Kimi represent different model capabilities and cost structures; Selection bias: Single expert baseline may not be representative of typical Pegasus developer practices; Comparison bias: Claude Code is tested with structured pegasus-ai plugin skills, while Codex and Kimi only have static documentation context—unequal conditions for comparison; Use case specificity: Evaluation limited to federated learning workflows; generalizability to other scientific workflow types unclear; LLM version differences: Different LLM versions used (Claude Opus 4, GPT-5.4, Kimi K2.6) with different training dates and capabilities; Small-scale data: TCIA and NIH chest X-ray datasets may not reflect production federated learning scenarios; Selection bias: federated learning use case may not represent all scientific workflow types; Evaluation limited to specific AI models (Claude, Codex, Kimi); results may not generalize to other LLMs; Manual expert baseline developed by one expert; may reflect individual preferences rather than general best practices; No blinded evaluation of generated vs. manual workflows
Limitations
- The authors explicitly state: "A quantitative evaluation of the MCP layer, including gateway routing behavior, multi-node coordination, and remote-client interaction patterns, is outside the scope of this paper." Additionally, the federated learning performance gap is attributed to data constraints: "the effective data available per client in the federated setting is limited
- For example, in the TCIA setup with K=10 clients, each client receives approximately 330 samples, which constrains local training capacity." The manual workflow baseline represents a single expert implementation, and the study does not systematically evaluate the MCP-based remote management layer quantitatively.
Open questions raised
- Future work will focus on evaluating the performance of the MCP layer, and building an AI-driven platform that guides users in designing more complex workflows, and manages the full workflow lifecycle - from specification to optimization of workflow runs across heterogeneous infrastructures.
- Quantitative evaluation of the MCP (Model Context Protocol) layer, including gateway routing behavior, multi-node coordination, and remote-client interaction patterns
- Performance optimization of workflow runs across heterogeneous infrastructures
- Extension to workflow types beyond federated learning (generalizability of the approach)
- Exploration of advanced federated learning design choices such as early stopping and alternative orchestration strategies (e.g., Ensemble Manager)
- Integration of federated learning-specific optimizations into the AI-assisted generation system
Explore related topics
Related papers
- Autonomous chemical research with large language modelsDaniil A. Boiko · 2023 · 809 citations
- To use or not to use ChatGPT in higher education? A study of students’ acceptance and use of technologyArtur Strzelecki · 2023 · 691 citations
- “ChatGPT seems too good to be true”: College students’ use and perceptions of generative AIClare Baek · 2024 · 98 citations
- AI chatbots in programming education: Students’ use in a scientific computing course and consequences for learningS.E.A. Groothuijsen · 2024 · 65 citations
- Language agents achieve superhuman synthesis of scientific knowledgeMichael Skarlinski · 2024 · 40 citations
- A Chain-of-Thought Prompting Approach with LLMs for Evaluating Students’ Formative Assessment Responses in ScienceClayton Cohn · 2024 · 38 citations