5 Sources
[1]
Anthropic AI 'formalizes' proof of Fermat's last theorem in just 11 days
Fermat's last theorem, one of the most celebrated mathematical results of the last half-century, has been turned into computer-verified code for the first time, using an advanced prototype of the artificial-intelligence (AI) chatbot Claude. The fact that a machine could turn the work of human
[2]
Fermat's last theorem formalised by AI agents in just 11 days | New Scientist
AI company Anthropic has created a formalised proof of Fermat's last theorem. It took just 11 days for a group of AI agents to complete the task, confirming that the human-found proof proposed in the 1990s is correct. Fermat's last theorem puzzled mathematicians for centuries until it was proven
[3]
Claude formalised Fermat's Last Theorem in 11 days
Dozens of Claude agents wrote 13 million lines of Lean in 11 days to produce the first computer-checked proof of Fermat's Last Theorem. The mathematician funded to do the same work says it tells us nothing about mathematics, and everything about what formalisation can now do. A mathematician holds
[4]
AI Just Solved a 350-Year-Old Math Problem By Writing the Longest Proof Ever
Kevin Buzzard, the mathematician leading that human project, reviewed Claude's proof and confirmed it holds up using nothing but math's most basic logical rules. Anthropic says its Claude AI just wrote the longest math proof ever made, and used it to formally prove Fermat's Last Theorem, a problem
[5]
Anthropic uses Claude to formalize proof of Fermat's Last Theorem
Anthropic uses Claude to formalize proof of Fermat's Last Theorem Anthropic PBC has used Claude to create a computer-verifiable version of a famous, highly complicated mathematical proof. The company detailed the project in a blog post published today. A proof is a series of arguments that
Share
Copy Link
Anthropic AI achieved a mathematical milestone by formalizing Fermat's Last Theorem in 11 days using Claude agents. The AI produced 13 million lines of computer-verified code and proved 29,500 intermediate theorems, completing a project that mathematicians expected would take a decade.
Anthropic AI has accomplished what mathematicians expected would take 10 years, completing a formalized proof of Fermat's Last Theorem in just 11 days
1
. Using its advanced Claude model, Anthropic PBC produced 13 million lines of computer-verified code in the Lean programming language, creating the largest formalized proof ever written2
. This achievement marks a pivotal moment in AI-aided mathematical reasoning, demonstrating that machines can now tackle the most complex mathematical work with minimal human oversight.The formalized proof translates Andrew Wiles' celebrated 1995 proof into computer-checkable format, eliminating any possibility of human error. Kevin Buzzard, a mathematician at Imperial College London who leads a five-year project funded with £1 million to accomplish the same task, confirmed the proof's validity, stating it relies on "no assumptions other than the axioms of mathematics"
3
. The AI formalizes proof that number theorist Alex Kontorovich at Rutgers University described as mind-blowing, fundamentally changing expectations about what autoformalization in mathematics can achieve1
.
Source: Decrypt
Fermat's Last Theorem states that no three positive whole numbers can satisfy the equation aⁿ + bⁿ = cⁿ when n is greater than 2
2
. French mathematician Pierre de Fermat posed this puzzle in 1637, famously claiming he had discovered a proof too large to fit in the margin of his textbook4
. For 358 years, mathematicians attempted to reconstruct what Fermat might have proven, making it one of mathematics' most celebrated unsolved problems.
Source: New Scientist
Andrew Wiles finally cracked the problem after seven years of secret work, announcing his breakthrough in 1993
2
. However, reviewers discovered a flaw that required Wiles and collaborator Richard Taylor nearly a year to fix before publishing the corrected 129-page proof in May 19954
. The proof earned Wiles an Abel Prize in 2016, and while solving this particular equation has limited practical application, the techniques developed helped unite distant mathematical disciplines1
.Dozens of Claude agents worked in parallel using an internal research model roughly comparable to Claude Fable 5.1, consuming approximately six billion output tokens during the 11-day run
3
. At standard pricing of $50 per million output tokens, this would translate to roughly $300,000 in computational costs3
. The AI agents proved 30,300 intermediate theorems, using 29,500 of them to build the complete computer-verified code5
.Human experts provided only minimal high-level instructions throughout the process. Researcher Tianyi Peng's guidance consisted of brief fragments like "Jacobian as a scheme sounds high priority" and "Push Mazur to be done soon"
3
. This limited human input demonstrates the autonomous capabilities of modern AI agents in handling complex mathematical work, a stark contrast to traditional collaborative mathematical research that requires constant human oversight and decision-making.Anthropic's initial attempt to formalize Andrew Wiles' proof failed when agents lost track of the project's state and stopped collaborating effectively
2
. The breakthrough came after implementing Prove2Me, an open-source coordination platform built by Tianyi Peng with collaborators at Columbia University3
. This tool maintains a graph of theorem statements, enabling AI agents to identify optimal next steps and avoid duplicating work.Prove2Me splits statements and proofs into separate files to accelerate compilation and keeps plain-language descriptions of each statement so work can be found and reused
3
. The failed early runs still contributed approximately 7% of the non-boilerplate lines in the final proof3
. This demonstrates that the scaffolding around a model—not just the model itself—determines what it achieves, with the same agents and weights producing failure then success based solely on coordination infrastructure.Related Stories
The 13-million-line formalized proof is more than five times the size of Mathlib, the community's central repository of formalized mathematics that currently contains 2 million lines
2
. Kevin Buzzard measured the proof at 13.4 million lines, noting it takes nearly 20 times as long to compile on a 96-core machine compared to existing Mathlib content3
. Daniel Litt, a number theorist at the University of Toronto, emphasized the significance: "If they can formalize Fermat's last theorem, they can probably formalize anything"1
.
Source: Nature
This achievement suggests AI could soon scrutinize the entire library of mathematical knowledge, potentially identifying errors in well-known results. Buzzard noted that such a scenario was "a fantasy" just two years ago
1
. The technology's ability to autoformalize algebra, harmonic analysis, geometry and number theory demonstrates that AI autoformalization artifacts are now robust enough to build upon, with multi-layered proofs becoming feasible2
.While the formalized proof changes nothing about the underlying mathematics—Buzzard estimates the mathematical community already placed confidence in Wiles' proof at 99.9% to 100%—it fundamentally transforms what's possible in research formalization
3
. If an AI swarm can formalize thousands of pages of literature end-to-end in 11 days, modern research formalization could begin happening in real-time alongside new discoveries. This represents a shift from formalization as a years-long verification exercise to a near-instantaneous validation tool.However, integration challenges remain. None of the 13 million lines can enter Mathlib in their current state, as the proof would need significant refactoring to meet community standards
3
. Buzzard also investigated whether the AI might have exploited known soundness bugs in Lean to artificially complete the proof, but found the code "plainly developing the mathematics the proof needs" without shortcuts3
. The milestone follows Anthropic's recent work using Claude to discover new information about the Riemann zeta function, while competitor OpenAI has deployed its Astra model to solve Erdos problems and advance theoretical computer science5
.Summarized by
Navi
[3]
21 Jul 2026•Science and Research

21 May 2026•Science and Research

01 Jul 2026•Science and Research

1
Policy and Regulation

2
Technology

3
Technology
