Dr.Sai: An agentic AI for real-world physics analysis at BESIII
Mingfeng He, Jiang Fayu, Junkun Jiao, M. H. Li, Ke Li, Yipu Liao et al. · ArXiv.org · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
Methodology & findings
Study design
Multi-agent system architecture with LLM-based autonomous workflow orchestration, validated through Monte Carlo simulation re-measurements of J/ψ branching fractions across ten decay channels.
Main result
Dr.Sai successfully automated the complete BESIII analysis chain for branching fraction measurements. The system "successfully managed the complete chain of physics analysis, from event selection and kinematic fitting to preliminary systematic uncertainty estimation, producing results in excellent agreement with Monte Carlo simulations and established physical benchmarks." Across ten J/ψ decay channels, the measured branching fractions showed high consistency with input values, demonstrating the system's capability for reproducible physics analysis.
Research paradigm
Computational empiricism with positivist assumptions about automated scientific discovery
Author conclusions
"This work establishes Dr.Sai as a viable framework for deploying multi-agent systems to accelerate scientific discovery in complex, real-world experimental environments." The authors note that "while reflection mechanisms enhance reliability, the primary bottlenecks for such autonomous systems remain the precise synthesis and invocation of scientific tools/code and the structural representation of domain expertise." They conclude that "The framework and principles demonstrated by Dr.Sai are also relevant to other data-intensive fields—such as astronomy and genomics—where the automation of complex analysis remains a fundamental challenge."
Risk of bias
Model selection bias: comparison limited to five specific LLM implementations; Task design bias: benchmark tasks selected may not represent full complexity spectrum of HEP analysis; Evaluation bias: success rate metrics may favor certain model architectures; Data bias: exclusive use of BESIII Monte Carlo samples limits generalizability; Model selection bias: Only frontier LLM models tested; limited diversity in base models; Evaluation bias: Success metrics based on task completion rather than physics accuracy validation; Data bias: Training uses only BESIII-specific knowledge; generalization to other experiments unclear; Confirmation bias: Validation limited to Monte Carlo samples; no independent real experimental data tested; Survivorship bias: Analysis focuses on successful task completions; failure cascades may propagate undetected; Model selection bias: frontier models (Qwen3-max, DeepSeek-v3.2, GLM-4.7) show significantly higher performance than others, potentially skewing generalizability; Limited experimental scope: validation performed only on MC samples equivalent to 2009 ψ(2S) dataset, not real experimental data; Systematic uncertainty limitations: incomplete framework focusing only on primary contributors; Task exclusion bias: complex decay channels deliberately excluded from analysis
Limitations
- "Due to the inherent complexity of systematic uncertainty estimation, the current implementation focuses on primary contributors
- In this study, we consider uncertainties arising from reconstruction efficiencies and cited input values
- A more comprehensive systematic framework remains a focus for future development." Additionally, "Channels containing an additional π+π− pair in the final state were excluded to avoid high mis-combination rates with the transition pions from ψ(2S), a complexity currently beyond the system's logic."
Open questions raised
- Comprehensive systematic uncertainty estimation framework needed beyond primary contributors
- Handling of complex decay channels with high mis-combination rates
- Refinement of tool synthesis precision and code invocation mechanisms
- Improved structural representation of domain expertise in LLM contexts
- Extension to other data-intensive scientific fields (astronomy, genomics)
- Comprehensive systematic uncertainty estimation framework needed
Explore related topics
Related papers
- Rayyan—a web and mobile app for systematic reviewsMourad Ouzzani · 2016 · 24,664 citations
- What if the devil is my guardian angel: ChatGPT as a case study of using chatbots in educationAhmed Tlili · 2023 · 1,587 citations
- Autonomous chemical research with large language modelsDaniil A. Boiko · 2023 · 809 citations
- Embracing the future of Artificial Intelligence in the classroom: the relevance of AI literacy, prompt engineering, and critical thinking in modern educationYoshija Walter · 2024 · 805 citations
- Leveraging ChatGPT for Enhancing Critical Thinking SkillsYing Guo · 2023 · 223 citations
- Students’ use of large language models in engineering education: A case study on technology acceptance, perceptions, efficacy, and detection chancesMargherita Bernabei · 2023 · 150 citations