AI Solves Math Problems

An 80-year-old conjecture, disproved by a machine

On 20 May 2026, OpenAI announced that one of its reasoning models had produced an original proof: a counterexample to a conjecture Paul Erdős posed in 1946. Nine mathematicians signed a commentary confirming it.

This was not a search of the literature. It was new mathematics, and it changes what you can reasonably expect an AI system to contribute to hard technical work.

Overview

Paul Erdős asked in 1946 how many pairs of points in a plane can be exactly one unit apart. The received wisdom was that the best arrangements look roughly like square grids. An internal, unreleased OpenAI reasoning model found an entirely new family of constructions that does better, disproving the conjecture. It reached the result with techniques from algebraic number theory, a field that had not previously been applied to this problem.

The model was not built for mathematics. It received the problem statement, got no hints about which approach to take, and was not guided step by step. It returned a complete proof of roughly 125 pages. OpenAI published it together with commentary signed by prominent mathematicians, among them Noga Alon, Melanie Wood and Thomas Bloom, who maintains the reference site for open Erdős problems.

This is worth stating plainly because the same company had overclaimed before. Roughly seven months earlier, an OpenAI executive said GPT-5 had solved ten Erdős problems. It had in fact located solutions that already existed in the literature. Bloom called that claim a dramatic misrepresentation. The May 2026 result is a different kind of thing, and the careful independent review is the reason we can say so.

How quickly this happened

December 2025A false start

Problem 333, retracted

An AI-assisted solution to Erdős problem 333 is announced and then withdrawn: Erdős himself had solved it in 1977. The episode sets the pattern for the whole year, in which the hard part is not producing an argument but checking whether it is new and whether it is correct.

January 2026First real results

Undergraduates and a DeepMind team

Kevin Barreto and Liam Price solve problem 728 with GPT-5.2 Pro, verifying consistency with Harmonic's Aristotle before Nat Sothanaphan formalises it. In parallel, a 24-researcher Google DeepMind team evaluates around 700 conjectures with Gemini, solving four and rediscovering solutions to nine more.

May 2026The breakthrough

The unit distance conjecture falls

A second DeepMind team autonomously resolves nine of 353 formalised open problems at a few hundred dollars per problem. Then, on 20 May, OpenAI announces the counterexample to Erdős's 1946 unit distance conjecture - by its own account the first time AI has autonomously solved a prominent open problem central to a field of mathematics.

1 August 2026It was not a one-off

Astra adds ten more advances

OpenAI reports that an unreleased model called Astra has produced ten further mathematical advances, including solutions to three more Erdős problems. Whatever the May result was, it was not luck.

What mathematicians actually said

The proof is serious work

Tim Gowers judged that the result approached the publication standard of an elite journal. Jacob Tsimerman described its sophistication as intimidating. This is not faint praise from people with no incentive to give it.

Verification is now the bottleneck

Thomas Bloom's warning is the most useful sentence in the whole debate: AI is being used heavily by people who are not mathematicians and cannot check the output, producing papers of 100 to 200 pages that no human has read. Generating an argument is cheap. Establishing that it is right is not.

Not everyone is delighted

Noga Alon of Princeton stopped working on Erdős problems: once AI started solving them, he said, there is no point any more. Terence Tao withdrew from the Erdős community to concentrate on his own research. A field can be advanced and demoralised at the same time.

The wins are clustered, not universal

The solved problems sit overwhelmingly in number theory, combinatorics and graph theory. That distribution says more about what current models are good at than about mathematics as a whole. Read the results as a sharp capability in a narrow band, not as general mathematical competence.

Why this matters outside mathematics

It is tempting to file this under interesting but irrelevant. That would be a mistake, because the result answers a question that comes up in every serious AI conversation: can these systems produce genuinely new work, or do they only recombine what they have seen?

The answer is now documented. A general-purpose model, not a specialist tool, solved a problem that strong human mathematicians had failed to solve for eighty years, using a technique nobody had thought to apply. If that is possible in mathematics, where the problems are unusually hard and the verification unusually strict, it is worth asking what it is possible in your own domain.

The concrete parallels are close at hand. Constrained optimisation, scheduling, tolerance analysis, material selection, tariff and pricing structures, test-case generation for a safety-critical component: these are all problems where the search space is large, the good answer is not obvious, and a company has usually settled on the arrangement that seemed reasonable years ago. That is exactly the shape of the unit distance problem, where everyone assumed a grid was optimal because a grid looked right.

And the caveat transfers as well. OpenAI's result stood up because nine mathematicians reviewed it before anyone celebrated. The same discipline applies to an AI-generated production plan or a rewritten pricing model: the value of the output is capped by your ability to verify it. Companies that get real returns from AI on hard technical questions are the ones that build the review step in from the start, rather than the ones with the best prompts.

Which of your assumptions has never been tested?

In a first conversation we go through the technical decisions in your company that were made once and never revisited - and assess honestly where a modern reasoning model can contribute something and where it cannot.

Arrange a conversation

Your first step to AI success

Your advisor, Ilirjan Bytyqi

Your contact

Ilirjan Bytyqi, M.Sc.Operations Manager at Ziya GmbH
Write to us
info@ziya.de