QED: An Open-Source Multi-Agent System for Generating Mathematical Proofs on Open Problems
Chenyang An, Qihao Ye, M. Pan, Jiayaun Zhang · ArXiv.org · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
Methodology & findings
Study design
Case study evaluation on five open research problems contributed by domain experts.
Main result
QED produced expert-verified original proofs for three of five open research problems in applied analysis and PDEs. As the authors state: "QED produces correct proofs for three of the five problems. Each proof was independently verified by the contributing domain experts, who confirmed the results to be original (not previously known) and nontrivial (requiring genuine mathematical insight)." The system demonstrated capability to generate original mathematical proofs without prior knowledge of solutions, including one proof that revealed an unexpected mathematical equivalence (Problem 4's Batchelor scale equivalence).
Research paradigm
Empiricist (system design driven by observed failure modes; validation through case studies)
Author conclusions
The authors conclude: "These results demonstrate that, with careful system design addressing known LLM weaknesses, AI can produce original, nontrivial proofs for open mathematical research." They further state: "These results suggest a practical role for AI in mathematical research—not as a replacement for human mathematicians but as a tool that can provide nontrivial candidate proofs." The expert assessment indicates "in his field, knowing the correct answer is of primary importance, and that these results indicate AI can serve as a useful tool for exploring conjectures."
Risk of bias
Selection bias: Only five open problems selected; may not be representative of broader research-level problems; Expert verification bias: Domain experts contributed the problems; their familiarity may introduce unconscious bias in assessment; Model selection bias: Different models used for different problem sets (GPT-5.4 for P1-4, Claude Opus for P5); Publication bias risk: Successful cases reported; unsuccessful mechanisms may remain undiscovered; Selection bias: Only five problems evaluated; problem selection criteria not specified as random or systematic; Confirmation bias: Domain experts who contributed problems may have favorable disposition toward AI results; Attrition bias: No discussion of how problems were selected or whether there were rejected problems before final five; Model selection bias: Different models used for different problem sets without clear justification; Publication bias: Only successful/interesting cases appear to be reported in detail; Selection bias: Only five open problems selected, all in applied analysis/PDEs domain; Domain expertise bias: Problems contributed by domain experts with specific research interests; Model selection bias: Different models used for different problem sets (GPT-5.4 vs Claude Opus 4.6); Verification bias: Expert verification may be influenced by knowledge of AI-generated nature of proofs; Publication bias: Only successful/interesting cases fully reported; negative cases relegated to appendices
Limitations
- "No benchmark exists for open research problems
- By definition, these problems lack known solutions against which to measure performance." Additionally, "verifying the correctness and originality of a proof requires substantial effort from domain experts
- Running QED with each component disabled would multiply the number of proofs requiring expert review, imposing an impractical burden on the mathematicians who volunteered their time." The system failed on two problems (P1 and P2), and the paper notes: "The expert noted that, as with Problem 4, the system does not attempt to closely follow the methodology of referenced papers, even when those papers are explicitly mentioned, and instead tries to derive a solution independently."
Open questions raised
- Extension to other mathematical domains beyond applied analysis and PDEs
- Evaluation on larger sets of open problems
- Development of benchmarks for open research-level proof tasks
- Investigation of why certain failure modes (e.g., misallocation of proof effort) persist despite architectural interventions
- Regarding publication potential: "the expert noted that while the two proofs are individually of interest, they would need to be supplemented with additional similar results to constitute a publishable paper"
- The paper identifies the gap that "producing genuinely novel, correct proofs for open problems with AI systems is still an outstanding challenge." It further notes that prior work (Google's Aletheia agent) showed only 13 valid solutions from 212 candidates. The authors frame their contribution as addressing this gap through systematic characterization of failure modes and architectural interventions.
Explore related topics
Related papers
- Autonomous chemical research with large language modelsDaniil A. Boiko · 2023 · 809 citations
- To use or not to use ChatGPT in higher education? A study of students’ acceptance and use of technologyArtur Strzelecki · 2023 · 691 citations
- “ChatGPT seems too good to be true”: College students’ use and perceptions of generative AIClare Baek · 2024 · 98 citations
- AI chatbots in programming education: Students’ use in a scientific computing course and consequences for learningS.E.A. Groothuijsen · 2024 · 65 citations
- Language agents achieve superhuman synthesis of scientific knowledgeMichael Skarlinski · 2024 · 40 citations
- A Chain-of-Thought Prompting Approach with LLMs for Evaluating Students’ Formative Assessment Responses in ScienceClayton Cohn · 2024 · 38 citations