12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

QED: An Open-Source Multi-Agent System for Generating Mathematical Proofs on Open Problems

Chenyang An, Qihao Ye, M. Pan, Jiayaun Zhang · ArXiv.org · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
1/4
Quality (LMQS)
D
Evidence
0
Citations
0.00
FWCI

Methodology & findings

Study design

Case study evaluation on five open research problems contributed by domain experts.

Main result

QED produced expert-verified original proofs for three of five open research problems in applied analysis and PDEs. As the authors state: "QED produces correct proofs for three of the five problems. Each proof was independently verified by the contributing domain experts, who confirmed the results to be original (not previously known) and nontrivial (requiring genuine mathematical insight)." The system demonstrated capability to generate original mathematical proofs without prior knowledge of solutions, including one proof that revealed an unexpected mathematical equivalence (Problem 4's Batchelor scale equivalence).

Research paradigm

Empiricist (system design driven by observed failure modes; validation through case studies)

Author conclusions

The authors conclude: "These results demonstrate that, with careful system design addressing known LLM weaknesses, AI can produce original, nontrivial proofs for open mathematical research." They further state: "These results suggest a practical role for AI in mathematical research—not as a replacement for human mathematicians but as a tool that can provide nontrivial candidate proofs." The expert assessment indicates "in his field, knowing the correct answer is of primary importance, and that these results indicate AI can serve as a useful tool for exploring conjectures."

Risk of bias

Selection bias: Only five open problems selected; may not be representative of broader research-level problems; Expert verification bias: Domain experts contributed the problems; their familiarity may introduce unconscious bias in assessment; Model selection bias: Different models used for different problem sets (GPT-5.4 for P1-4, Claude Opus for P5); Publication bias risk: Successful cases reported; unsuccessful mechanisms may remain undiscovered; Selection bias: Only five problems evaluated; problem selection criteria not specified as random or systematic; Confirmation bias: Domain experts who contributed problems may have favorable disposition toward AI results; Attrition bias: No discussion of how problems were selected or whether there were rejected problems before final five; Model selection bias: Different models used for different problem sets without clear justification; Publication bias: Only successful/interesting cases appear to be reported in detail; Selection bias: Only five open problems selected, all in applied analysis/PDEs domain; Domain expertise bias: Problems contributed by domain experts with specific research interests; Model selection bias: Different models used for different problem sets (GPT-5.4 vs Claude Opus 4.6); Verification bias: Expert verification may be influenced by knowledge of AI-generated nature of proofs; Publication bias: Only successful/interesting cases fully reported; negative cases relegated to appendices

Limitations

  • "No benchmark exists for open research problems
  • By definition, these problems lack known solutions against which to measure performance." Additionally, "verifying the correctness and originality of a proof requires substantial effort from domain experts
  • Running QED with each component disabled would multiply the number of proofs requiring expert review, imposing an impractical burden on the mathematicians who volunteered their time." The system failed on two problems (P1 and P2), and the paper notes: "The expert noted that, as with Problem 4, the system does not attempt to closely follow the methodology of referenced papers, even when those papers are explicitly mentioned, and instead tries to derive a solution independently."

Open questions raised

  • Extension to other mathematical domains beyond applied analysis and PDEs
  • Evaluation on larger sets of open problems
  • Development of benchmarks for open research-level proof tasks
  • Investigation of why certain failure modes (e.g., misallocation of proof effort) persist despite architectural interventions
  • Regarding publication potential: "the expert noted that while the two proofs are individually of interest, they would need to be supplemented with additional similar results to constitute a publishable paper"
  • The paper identifies the gap that "producing genuinely novel, correct proofs for open problems with AI systems is still an outstanding challenge." It further notes that prior work (Google's Aletheia agent) showed only 13 valid solutions from 212 candidates. The authors frame their contribution as addressing this gap through systematic characterization of failure modes and architectural interventions.
Data: Problem statements are provided in Appendices B and C for negative cases. Problem 5 statement and proof are stated as withheld pending ArXiv posting of corresponding mathematics paper. Full proofs for Problems 3-4 referenced as in Appendix A.Code: QED is released as open-source software (specific repository URL not provided in paper); QED is described as "an open-source multi-agent proof system" and "released as open-source software," but no specific repository URL is provided in the paper.; QED is released as open-source software (specific repository URL not provided in document)Extracted from: pdfAgreement 58%

Explore related topics

Related papers