From Experimental Limits to Physical Insight: A Retrieval-Augmented Multi-Agent Framework for Interpreting Searches Beyond the Standard Model
A. Çakır, Ayca Yerlikaya · ArXiv.org · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
Methodology & findings
Study design
Design Science with case studies.
Primary method
Design Science approach with multi-agent AI system architecture design
Main result
The experimental results demonstrate that the proposed system is capable of retrieving experimental measurements, reconstructing physics plots, and synthesizing cross-paper insights. "By reconstructing exclusion limits from structured HEPData datasets, the system enables quantitative analysis and comparison of collider search constraints reported across multiple publications." The case studies show that HEP-CoPilot can retrieve relevant measurements, reconstruct exclusion limits directly from HEPData records, and perform cross-paper comparisons of experimental constraints.
Research paradigm
Design Science / Pragmatist
Author conclusions
"These results suggest that retrieval-augmented, domain-aware AI systems can provide an effective interface for navigating the increasingly complex landscape of particle physics literature. By assisting researchers in retrieving relevant experimental evidence, reconstructing numerical results, and synthesizing insights across multiple publications, systems such as HEP-CoPilot have the potential to function as scientific co-pilots that augment the reasoning process of physicists working with large bodies of experimental data." The authors conclude that "as experimental results and scientific publications continue to grow in volume and complexity, tools that facilitate structured exploration of the literature may become increasingly important for accelerating the interpretation of new physics searches."
Risk of bias
Limited sample of CMS analyses (only 3 publications due to computational constraints); Evaluation bias from single model assessment mitigated by using multiple independent LLM judges; Dependence on HEPData availability may bias toward well-documented analyses; Local model deployment may have capability limitations affecting reasoning quality; Selection bias: only three CMS analyses evaluated due to computational constraints; Limited representativeness: restricted to recent CMS searches, not ATLAS or other experiments; Evaluation bias: LLM-as-a-Judge methodology may reflect biases in the three judge models selected; No ground-truth comparison: baseline is PDF-based querying without independent validation; Limited evaluation dataset (only 3 CMS analyses); Dependence on LLM judges potentially biased toward certain model outputs; Local language model capabilities may be limited compared to proprietary systems; Selection of experimental analyses may not represent broader particle physics literature
Limitations
- "The current study evaluates the system using a limited number of CMS analyses
- This restriction is primarily due to computational constraints, as the prototype implementation was developed and tested on a local computing environment with limited storage and indexing capacity
- Consequently, only three publications were incorporated into the experimental evaluation." Additionally, "the current implementation employs locally deployed language models through the Ollama framework
- While this approach provides flexibility and enables offline experimentation, the reasoning capabilities of local models may be more limited compared to larger state-of-the-art systems such as ChatGPT or Claude." The authors also note that "some information remains embedded in figures or complex tables within the papers themselves" and that "as with all language-model-based systems, ensuring reliability and minimizing the risk of hallucinated interpretations remains an important challenge."
Open questions raised
- Extension to larger collections of particle physics publications using higher-capacity computing infrastructure
- Integration of more advanced language models beyond locally deployed systems
- Incorporation of automated figure digitization and table extraction techniques
- Development of physics-aware reasoning modules for theoretical interpretation and cross-section calculations
- Improvement of verification mechanisms and uncertainty-aware reasoning to minimize hallucinated interpretations
- Limited scalability: current system restricted to three publications due to computational constraints; future work should deploy on higher-capacity infrastructure for larger publication collections
Explore related topics
Related papers
- Rayyan—a web and mobile app for systematic reviewsMourad Ouzzani · 2016 · 24,664 citations
- Guidance for conducting systematic scoping reviewsMicah D.J. Peters · 2015 · 7,472 citations
- Which academic search systems are suitable for systematic reviews or meta‐analyses? Evaluating retrieval qualities of Google Scholar, PubMed, and 26 other resourcesMichael Gusenbauer · 2019 · 2,116 citations
- Hallucination Rates and Reference Accuracy of ChatGPT and Bard for Systematic Reviews: Comparative AnalysisMikaël Chelli · 2024 · 295 citations
- Artificial intelligence and the conduct of literature reviewsGerit Wagner · 2021 · 275 citations
- Artificial intelligence for literature reviews: opportunities and challengesF. J. Bolaños · 2024 · 188 citations