Epoch: 68 open problems, a Lean proof, and almost nobody clears it
On 1 September 2026, Epoch AI publishes FrontierMath Erdős. 68 problems still open in August 2026, formalized in Lean. On the first run, only a pre-release GPT-6 Astra solves 2. The other four models tested score zero.

Original AI-generated illustration · SecuFocus
At a glance
Key points
This is not a chat ranking. It is a math benchmark with a fixed budget and a verified proof. We did not rerun the models.
What changes
Epoch’s 1 September 2026 note describes 68 Erdős problems, still open in August 2026, selected with Thomas Bloom and written in Lean. The model must prove or disprove them inside a fixed budget. A solution counts only if the Lean proof passes verification.
The first run covers five models: a pre-release GPT-6 Astra, GPT-5.6 Sol, GPT-5.5, Claude Fable 5.1 and Claude Fable 5. One attempt per problem, 300 dollars and 72 hours. Astra scores 3 percent, 2 of 68. The others score 0 percent.
What leaves
The problems go to the model provider, along with the agent’s working time. Epoch prices two Astra successes: problem 74 disproved by a counterexample, 222 dollars and 10 hours; problem 126 proved, 172 dollars and 10 hours. The rest of the run stops when the budget runs out, with no verified proof.
What we did not check
We did not rerun the benchmark or open the Lean repository. Astra here is a pre-release, not the model in your chat. Epoch also says it made further attempts, outside the protocol, with larger budgets. Those attempts are not the score.
The choice
A 3 percent score on open problems is not a homework grade. It is a sign that a hard, verified benchmark still separates models. If a lab cites “Erdős” without the protocol, the budget and the Lean proof, it is not this benchmark.
Check and explore
Sources for this article
Numbers connect each reference to the passages that use it. Dates show when the documentation was consulted.
This article draws on the sources above. The exercises are for you to try on your devices; SecuFocus does not present them as tests carried out by its editorial team. Interfaces and features can change. Method and corrections.
Cite this article
Keep this reference with the article when you save or share it.