OpenAI Unveils AI-Generated Solutions to 100+ Unsolved Math Problems, Raising Concerns Over Lack of Human Understanding
By admin | Oct 08, 2026 | 3 min read
OpenAI recently unveiled hundreds of purported solutions to some of the most challenging mathematical problems known to humanity. The organization stated that it had sought guidance from a select panel of top mathematicians to prevent the kind of backlash that occurred the last time one of its models cracked a longstanding mathematical puzzle. However, OpenAI's efforts appear to have fallen short of the standards that these experts emphasized—particularly regarding the necessity of human comprehension of mathematical outcomes.
This shortfall is especially troubling in light of a newly published paper that identified inconsistencies between the natural-language explanations and formally verified solutions for a million-dollar problem that OpenAI's models reportedly solved.
The Advisory Group on Mathematics and Artificial Intelligence (AGMAI), based at Princeton University's Institute for Advanced Studies, comprises nine distinguished researchers from institutions worldwide. The group issued recommendations for frontier labs tackling mathematical problems at the close of September. In a statement regarding OpenAI's latest batch of proofs, AGMAI noted that "it is ultimately up to the mathematical community to assess the extent to which our recommendations were followed successfully."
Yet the group's foremost recommendation was "to stop testing advanced mathematical problems on proprietary models." OpenAI's announcement explicitly states that it is assessing its proprietary models using open research problems in mathematics. The lab evidently adhered to certain principles, such as releasing results promptly and providing details on how the models arrived at their conclusions. But this was not done for all cases: only ten out of 719 manuscripts included the model's chain of thought. For papers that remain incomprehensible to people, the mathematicians recommended that proofs be formalized—yet only 42% of the proofs OpenAI released had not undergone this formalization process.
Fundamentally, it remains uncertain whether OpenAI is fulfilling its "responsibility for ensuring that human understanding will follow" when it releases its proofs, as outlined in the AGMAI principles. AGMAI proposed that OpenAI should contribute funding to support human mathematicians whose work is essential to making the lab's solutions genuinely meaningful. "Problems are being solved autonomously by AI prompters who have no interest in the broader field itself once their initial target is 'solved', and do not understand the AI output well enough to answer questions on the result, give talks, or otherwise interact with the rest of the field," wrote Terence Tao, a renowned mathematician who has been critical of OpenAI's methods, on social media following the release.
This issue is illustrated by a paper published this week by mathematicians from the University of Cambridge and King's College London, which scrutinizes how frontier labs are tackling these challenges. When AI models solve mathematical problems, they initially generate a "natural language" explanation, then attempt to translate that result into Lean, a programming language that theoretically verifies the proof's accuracy by compiling it as code. However, issues can arise in how models convert their natural language proofs into code; this paper documents at least two discrepancies between the natural language proof and the Lean code underlying the solution OpenAI provided for a problem derived from the Navier-Stokes equations, which describe the complex behavior of fluids. These discrepancies don't necessarily invalidate either solution, but they do prompt questions about whether we can depend solely on models to formalize their own solutions without human oversight.
That's one reason AGMAI urged OpenAI to "include machine-readable metadata correlating the natural language and formal artifacts," something the frontier lab failed to do with these releases. "Because of the phenomenon of mistranslations - as highlighted in this paper - the NL proof by OpenAI and other autoformalised Lean proofs should not prima facie be trusted without the same peer review process and scrutiny that other proofs are subjected to," the authors of the "lost in translation" paper conclude.
Mathematicians emphasize that when humans discover new results, they take ownership of them and engage with the wider community through papers, talks, and seminars. This process deepens understanding of the solutions, uncovers strategies applicable to other problems, and enables the new knowledge to be put to practical use.
Comments
Please log in to leave a comment.
No comments yet. Be the first to comment!