Towards Autonomous Mathematics Research
Tony Feng, Trieu Trinh, Garrett Bingham, Dawsen Hwang, Yuri Chervonyi, Junehyuk Jung et al. · ArXiv.org · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
Methodology & findings
Study design
The methodology comprises multiple empirical and case-study components: (1) Scaling law evaluation using the IMO-Proof Bench Advanced (30 problems) and internal FutureMath Basic PhD-level benchmark, with models evaluated at multiple compute scales (2^7 to 2^12); (2) Systematic deployment of Aletheia on 700 open Erdős problems from December 2-9, 2025, with human expert evaluation in three tiers (initial verification, narrowing to manageable pool, domain expert vetting); (3) Comparative ablation studies using Gemini Deep Think (IMO Gold scale) on identical prompts; (4) FirstProof benchmark evaluation (10 research-level problems) with pre-determined verification prompts and expert consensus evaluation; (5) Agent architecture design (Generator-Verifier-Reviser orchestration) with natural language verification and extensive tool use (Google Search, web browsing, Python execution).
Main result
The paper demonstrates that Aletheia, a math research agent, achieved "a 93% overall score without any tool usage" on IMO-Proof Bench Advanced, "surpassing Deep Think across all tested compute scales using the same base model." On the FutureMath Basic benchmark for PhD-level problems, "Aletheia again outperformed Deep Think at all compute scales on the same base model." Critically, the authors found that "on FirstProof problems, both of our runs on Aletheia produced solution candidates to exactly 6 problems (P2, P5, P7, P8, P9, P10)" with "all 6 problems were solved correctly (under the interpretation of being publishable after minor revisions)." On the broader Erdős problems benchmark, "out of the 200 solution candidates that we were definitely able to mark correct or incorrect, 137 (68.5%) of the responses were fundamentally flawed, while 63 (31.5%) of the responses were technically correct, of which only 13 (6.5%) were meaningfully correct."
Research paradigm
empirical-computational with formal mathematical reasoning
Author conclusions
The authors conclude: 'Ultimately, we believe that AI will become a tool that enhances rather than replaces mathematicians. Currently, natural language models struggle to reason reliably without human intervention to correct mistakes and hallucinations, while formal verification systems are not yet capable of even formulating the questions of interest on most research frontiers.' They further conclude: 'To date, hype notwithstanding, the impact of artificial intelligence on pure mathematics research has been limited. While our results do solve some problems that seem to have eluded experts, they do not indicate that artificial intelligence has matched, or will match, the capabilities of human mathematicians. Rather, they illustrate how certain comparative advantages of AI models over humans can be useful for certain kinds of problems.' The authors note: 'This perhaps clarifies the directions where human researchers can expect the most impact from AI in the near future. A first observation is that AI models exhibit a form of intelligence that diverges significantly from that of human scientists. In any specific subject, frontier models have much shallower knowledge than a domain expert, but they also possess superhuman breadth of knowledge, which could be the key to unlocking certain problems.'
Risk of bias
Selection bias: Success cases reported are non-representative; failure cases underreported on public channels; Evaluation bias: Human expert grading on novel problems lacks standardized rubrics; subjectivity in mathematical significance assessment; Data contamination risk: Knowledge cutoff of base models falls between IMO 2024 and 2025, with possible exposure to IMO 2024 problems; Confirmation bias: Authors may preferentially highlight positive AI contributions; mathematicians may overstate AI involvement for publicity; Attrition bias: Only 212 of 700 initial Erdős problem candidates returned solutions; high filtering rate reduces representativeness; Publication bias: Failures on FirstProof not comprehensively documented; baseline performance (2 problems solved out-of-box by GPT-5.2) may underrepresent competing systems; Selection bias in problem choice: success cases represent cherry-picked problems from wider benchmarking efforts where most problems showed no autonomous progress; Publication bias: only positive results documented; authors acknowledge 'Because only success cases tend to be reported in public forums, these results do not provide a complete picture of AI capability'; Evaluation subjectivity: mathematical significance assessment relies on expert judgment with potential disagreement (e.g., FirstProof Problem 8 had non-unanimous assessment: 5/7 experts rated as correct); Data contamination risk: model's knowledge cutoff postdates IMO 2024, raising questions about memorization versus reasoning capability; Confirmation bias in literature review: authors may have missed existing solutions due to search limitations when classifying Erdős problems as novel; Expert availability bias: assessment limited to available mathematicians, potentially missing domain specialists who could evaluate significance differently; Tool use confounding: Aletheia's performance improvements may partially reflect tool access (Google Search, web browsing) rather than core reasoning advances; Selection bias in choice of which problems to test (success cases reported more frequently than failures); Data contamination risk for IMO 2024 and 2025 problems (model knowledge cutoff post-dates competition dates); Researcher bias in classification of AI contributions to mathematics (subjective assessment of 'novelty' and 'significance'); Publication bias toward positive outcomes (negative results not comprehensively documented)
Limitations
- The authors identify several limitations: "To date, autonomous results have been relatively brief and elementary in comparison to typical human papers
- Furthermore, success cases seem to arise from clever technical manipulations or vast knowledge retrieval, rather than what mathematicians would consider to be genuine creativity, although the latter concept is admittedly subjective." Additionally, "Even with its verifier mechanism, Aletheia is still more prone to errors than human experts
- Furthermore, whenever there is room for ambiguity, the model exhibits a tendency to misinterpret the question in a way that is easiest to answer, even when such an interpretation would be obviously unintended to a human expert." The authors also note: "hallucination is still a common failure mode
- Even with internet search capability to check references, the model tends to fabricate or misrepresent results from legitimate references in order to assert a solution." Regarding Erdős problems specifically: "we made considerable efforts to review the literature, it is certainly possible that we missed earlier solutions to these problems by human mathematicians
- Therefore, our initial classification into categories is, at best, an upper bound on novelty."
Open questions raised
- The authors identify several gaps and future directions: (1) Inference-time scaling alone is insufficient for research-level mathematics; need for further improvements beyond scaling laws. (2) Tool use remains imperfect; integration of Python yielded only 'marginal improvements in mitigating computational hallucinations,' suggesting need for more specialized tools. (3) Autonomous results remain 'relatively brief and elementary in comparison to typical human papers'; success cases appear to arise from 'clever technical manipulations or vast knowledge retrieval, rather than what mathematicians would consider to be genuine creativity.' (4) Need for community-wide standards on documenting AI-assisted mathematics: 'For Level C or Level A results, where AI input is deemed essential, a possible baseline would be to expose at least the most important raw prompts and outputs that contain the essential new insights generated by AI. Specific standards are left for the community of mathematicians to decide.' (5) Formal verification systems remain incapable of formulating frontier research questions. (6) Understanding of when and why AI succeeds on certain problem types remains incomplete.
- Authors identify several gaps and future directions: (1) Broader application of inference-time scaling laws beyond competition mathematics to open-ended research problems; (2) Development of more sophisticated tools beyond standard code execution for computational verification; (3) Improvement of verifier mechanisms to better recognize hallucinations and citation errors; (4) Reduction of model tendency to misinterpret ambiguous problem statements through specification gaming; (5) Better handling of formal requirements (e.g., elementary-only proofs for IMO competition format); (6) Extension of natural language verification to match formal verification standards; (7) Community-wide standards for documenting and evaluating AI-assisted mathematics; (8) Further investigation of which problem types benefit most from AI's comparative advantages in breadth versus depth of knowledge.
- Broader standardization needed for evaluating and communicating AI-assisted mathematics research
- Need for formal verification systems capable of formulating frontier research questions
- Further development of AI agents to reduce hallucination and improve reliability without human intervention
- Systematic study of which problem classes benefit most from AI assistance
Explore related topics
Related papers
- Rayyan—a web and mobile app for systematic reviewsMourad Ouzzani · 2016 · 24,664 citations
- What if the devil is my guardian angel: ChatGPT as a case study of using chatbots in educationAhmed Tlili · 2023 · 1,587 citations
- Autonomous chemical research with large language modelsDaniil A. Boiko · 2023 · 809 citations
- Embracing the future of Artificial Intelligence in the classroom: the relevance of AI literacy, prompt engineering, and critical thinking in modern educationYoshija Walter · 2024 · 805 citations
- To use or not to use ChatGPT in higher education? A study of students’ acceptance and use of technologyArtur Strzelecki · 2023 · 691 citations
- Leveraging ChatGPT for Enhancing Critical Thinking SkillsYing Guo · 2023 · 223 citations