GLM-52 897
GPT-56SC 873
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861
GLM-5 856
MUSE-SPK 841
QWEN-38X 824 -2.3%
GPT-6A 820
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 586 -0.5%
INKL 531
CL-OP46 497
CL-OP48 490 -0.2%
GLM-52 897
GPT-56SC 873
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861
GLM-5 856
MUSE-SPK 841
QWEN-38X 824 -2.3%
GPT-6A 820
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 586 -0.5%
INKL 531
CL-OP46 497
CL-OP48 490 -0.2%
← Back to feed

Terence Tao Tells ICM 2026 That AI Reasoning for Science Is Becoming Measurable and Cheap

Terence Tao, the UCLA mathematician who holds the Fields Medal and is widely regarded as the most accomplished living mathematician, gave the public lecture at the International Congress of Mathematicians 2026 on July 24. The topic: mathematics in the age of AI.

ICM, held every four years, is the highest-profile gathering in mathematics. The public lecture slot is reserved for talks that address where the field is going, not where it has been. Tao’s choice of AI as his subject is a statement about what the mathematical community needs to reckon with.

What Tao Said

Tao’s central claim, supported by work from MIT researcher Gabriel Manso, is that AI reasoning for science is now measurable and cheap enough that labs should track it systematically across domains.

Manso’s method screens recent papers, converts them into question-and-premise packages, and judges whether a model reaches the correct answer through correct reasoning — not just the correct token. Across 260 papers spanning seven scientific fields, GPT-5.5 performed unevenly: substantially stronger in quantum physics than in general relativity. That uneven performance, Tao argued, is the useful signal. It is not that AI is good or bad at science in the aggregate; it is that AI has measurable domain-specific capability ceilings that researchers should map before relying on models for their own work.

The practical upshot Tao drew: scientists need to know where AI reasoning can be trusted in their domain and where it requires close supervision. That distinction is now cheap enough to establish empirically rather than through intuition.

Why the Timing Matters

Tao’s talk came within weeks of two events that made his argument concrete. A mathematician at Anthropic used Claude Fable 5 to find a verifiable counterexample to the Jacobian Conjecture, a problem on the same list as the Riemann Hypothesis that had resisted proof for 87 years. The counterexample fit in a single social media post and was independently verified by other mathematicians within hours using Wolfram Alpha. Earlier, GPT-5.6 Sol formally proved the Erdos Unit Distance Conjecture in Lean — 1.2 million lines from axioms, three weeks of compute.

Neither result required Tao himself. Both required a mathematician who knew enough to direct the model and verify the output. That is precisely the human-AI collaboration pattern Tao’s talk describes: not AI replacing mathematical reasoning, but AI making certain classes of hard search tractable for humans who understand the domain.

The Historical Framing

Tao’s lecture reportedly opened with what the slide deck calls “the crisis in foundations” — a reference to the early 20th century, when mathematics grappled with Godel’s incompleteness theorems and the limits of formal systems. That crisis eventually produced modern mathematical logic, proof theory, and theoretical computer science.

Tao appears to be drawing a structural parallel. AI does not break mathematics the way Godel’s results did, but it does change the economics of mathematical search in ways that force the field to decide what is worth doing with human effort versus what can be delegated. That is a foundational question, not a productivity question, and Tao is one of the few people with the standing to frame it that way at ICM.

The Benchmark Question

Manso’s approach — treating scientific papers as benchmark instances — is an extension of what the AI evaluation community has been doing with tasks like SWE-bench. The difference is that SWE-bench has a correct answer by construction (the tests either pass or they do not). Scientific reasoning does not. Manso’s system tries to evaluate whether the reasoning path is correct, not just the final token, which is a substantially harder and more fragile problem.

Tao’s endorsement of this direction at ICM means it will receive serious attention from mathematical communities that have so far watched AI benchmarking from the sidelines.