GPT-5.6 Sol Formally Proves Erdős Unit Distance Conjecture in Lean — 1.2 Million Lines From Axioms in 3 Weeks
Kevin Buzzard is a mathematician at Imperial College London, one of the core maintainers of mathlib — the comprehensive Lean mathematical library that took nine years of organized human effort to reach 2.3 million lines of code. On July 20 he published a detailed account of what he has watched happen to his field over the past seven weeks.
The short version: AI systems have now formally proved a major open conjecture from mathematical axioms, using machine-checkable Lean code, in a three-week run that generated more than half of what mathematicians spent nine years building.
Three Events, Seven Weeks
May 20, 2026. ChatGPT disproved the Erdős Unit Distance conjecture — a longstanding problem in discrete geometry asking about the maximum number of unit distances among a set of points in the plane. The proof structure used the Golod-Shafarevich theorem, a result in algebraic number theory from the 1960s. Human mathematicians with early access verified the argument. No Lean formalization existed.
May 26, 2026. Mike Freedman — Fields Medallist and Chief Science Officer of Logical Intelligence, a company cofounded with Turing Award winner Yann LeCun — emailed Buzzard to say his company had autoformalized the ChatGPT proof in Lean within days of its publication. The formalization covered the specific logical step: that Golod-Shafarevich implies the Erdős counterexample. Significant, but incomplete. The Golod-Shafarevich theorem itself — requiring more than 100 pages of global class field theory to prove — remained an assumption, not a formalized fact.
June 26, 2026. Boris Alexeev at OpenAI announced that he had steered GPT-5.6 Sol to a complete formalization of the Erdős proof from axioms — no remaining assumptions. Sol generated 1.2 million lines of Lean code over three weeks, including the hard theorems in global class field theory that Logical Intelligence had left unproven. Buzzard reviewed the code in a sandbox and confirmed the core claim: Sol had produced working proofs of nontrivial theorems in the cohomology of number fields.
Scale Reference
Mathlib contains 2.3 million lines of Lean code, built by an organized community of mathematicians over nine years. Sol produced 1.2 million lines in three weeks — without prior formalization infrastructure for the relevant number theory.
Global class field theory is the subfield involved. Buzzard had run a Clay Summer School in 2025 dedicated to formalizing it; the local case remains a PhD project. The global case, as of early 2026, he described as “a fantasy” to formalize. Sol did it as part of a three-week run on a single problem.
”Large AI-Generated Mathematics Is Inevitable”
Buzzard writes directly about the implication: “Perhaps it was at this point that the penny really dropped for me — large AI-generated developments of mathematics are inevitable.”
The standard workflow in formal mathematics — conjecture, informal proof, peer review, optional Lean formalization as a later community project — is being disrupted at every stage. The ChatGPT proof came first. Logical Intelligence formalized the top level in days. Sol formalized everything, including the hard foundational theorems, in weeks.
Pattern, Not Incident
Buzzard’s July 20 post does not stop at the Erdős counterexample. His account documents additional cases — including activity in algebraic group schemes — where AI systems are outpacing human counterexample-finders in domains where progress had stalled. The emerging pattern is AI finding and formally certifying counterexamples in fields where informal human consensus had treated the underlying questions as slow-moving or unsettled.
The Lean formalization angle matters here specifically. Informal proofs require trusting human reviewers. Machine-verified Lean code does not. When Sol generates 1.2 million lines and Buzzard can run it in a sandbox and confirm it proves what it claims to prove, the trust infrastructure changes. Whether the result is correct is no longer a question of authority — it is a question of execution.
That is new. And from Buzzard’s perspective, it arrived several years earlier than expected.