Anthropic AI Completes 10-Year Mathematical Proof Project in Just 11 Days Using Claude

Reviewed byNidhi Govil

5 Sources

Share

Anthropic AI achieved a mathematical milestone by formalizing Fermat's Last Theorem in 11 days using Claude agents. The AI produced 13 million lines of computer-verified code and proved 29,500 intermediate theorems, completing a project that mathematicians expected would take a decade.

Anthropic AI Achieves Mathematical Breakthrough with Fermat's Last Theorem

Anthropic AI has accomplished what mathematicians expected would take 10 years, completing a formalized proof of Fermat's Last Theorem in just 11 days

1

. Using its advanced Claude model, Anthropic PBC produced 13 million lines of computer-verified code in the Lean programming language, creating the largest formalized proof ever written

2

. This achievement marks a pivotal moment in AI-aided mathematical reasoning, demonstrating that machines can now tackle the most complex mathematical work with minimal human oversight.

The formalized proof translates Andrew Wiles' celebrated 1995 proof into computer-checkable format, eliminating any possibility of human error. Kevin Buzzard, a mathematician at Imperial College London who leads a five-year project funded with £1 million to accomplish the same task, confirmed the proof's validity, stating it relies on "no assumptions other than the axioms of mathematics"

3

. The AI formalizes proof that number theorist Alex Kontorovich at Rutgers University described as mind-blowing, fundamentally changing expectations about what autoformalization in mathematics can achieve

1

.

Source: Decrypt

Source: Decrypt

Understanding Fermat's Last Theorem and Its Historical Significance

Fermat's Last Theorem states that no three positive whole numbers can satisfy the equation aⁿ + bⁿ = cⁿ when n is greater than 2

2

. French mathematician Pierre de Fermat posed this puzzle in 1637, famously claiming he had discovered a proof too large to fit in the margin of his textbook

4

. For 358 years, mathematicians attempted to reconstruct what Fermat might have proven, making it one of mathematics' most celebrated unsolved problems.

Source: New Scientist

Source: New Scientist

Andrew Wiles finally cracked the problem after seven years of secret work, announcing his breakthrough in 1993

2

. However, reviewers discovered a flaw that required Wiles and collaborator Richard Taylor nearly a year to fix before publishing the corrected 129-page proof in May 1995

4

. The proof earned Wiles an Abel Prize in 2016, and while solving this particular equation has limited practical application, the techniques developed helped unite distant mathematical disciplines

1

.

How Claude's Multi-Agent System Completed the Formalized Proof

Dozens of Claude agents worked in parallel using an internal research model roughly comparable to Claude Fable 5.1, consuming approximately six billion output tokens during the 11-day run

3

. At standard pricing of $50 per million output tokens, this would translate to roughly $300,000 in computational costs

3

. The AI agents proved 30,300 intermediate theorems, using 29,500 of them to build the complete computer-verified code

5

.

Human experts provided only minimal high-level instructions throughout the process. Researcher Tianyi Peng's guidance consisted of brief fragments like "Jacobian as a scheme sounds high priority" and "Push Mazur to be done soon"

3

. This limited human input demonstrates the autonomous capabilities of modern AI agents in handling complex mathematical work, a stark contrast to traditional collaborative mathematical research that requires constant human oversight and decision-making.

The Critical Role of Prove2Me in Achieving Success

Anthropic's initial attempt to formalize Andrew Wiles' proof failed when agents lost track of the project's state and stopped collaborating effectively

2

. The breakthrough came after implementing Prove2Me, an open-source coordination platform built by Tianyi Peng with collaborators at Columbia University

3

. This tool maintains a graph of theorem statements, enabling AI agents to identify optimal next steps and avoid duplicating work.

Prove2Me splits statements and proofs into separate files to accelerate compilation and keeps plain-language descriptions of each statement so work can be found and reused

3

. The failed early runs still contributed approximately 7% of the non-boilerplate lines in the final proof

3

. This demonstrates that the scaffolding around a model—not just the model itself—determines what it achieves, with the same agents and weights producing failure then success based solely on coordination infrastructure.

Implications for Mathematical Research and Verification

The 13-million-line formalized proof is more than five times the size of Mathlib, the community's central repository of formalized mathematics that currently contains 2 million lines

2

. Kevin Buzzard measured the proof at 13.4 million lines, noting it takes nearly 20 times as long to compile on a 96-core machine compared to existing Mathlib content

3

. Daniel Litt, a number theorist at the University of Toronto, emphasized the significance: "If they can formalize Fermat's last theorem, they can probably formalize anything"

1

.

Source: Nature

Source: Nature

This achievement suggests AI could soon scrutinize the entire library of mathematical knowledge, potentially identifying errors in well-known results. Buzzard noted that such a scenario was "a fantasy" just two years ago

1

. The technology's ability to autoformalize algebra, harmonic analysis, geometry and number theory demonstrates that AI autoformalization artifacts are now robust enough to build upon, with multi-layered proofs becoming feasible

2

.

What This Means for the Future of Mathematical Collaboration

While the formalized proof changes nothing about the underlying mathematics—Buzzard estimates the mathematical community already placed confidence in Wiles' proof at 99.9% to 100%—it fundamentally transforms what's possible in research formalization

3

. If an AI swarm can formalize thousands of pages of literature end-to-end in 11 days, modern research formalization could begin happening in real-time alongside new discoveries. This represents a shift from formalization as a years-long verification exercise to a near-instantaneous validation tool.

However, integration challenges remain. None of the 13 million lines can enter Mathlib in their current state, as the proof would need significant refactoring to meet community standards

3

. Buzzard also investigated whether the AI might have exploited known soundness bugs in Lean to artificially complete the proof, but found the code "plainly developing the mathematics the proof needs" without shortcuts

3

. The milestone follows Anthropic's recent work using Claude to discover new information about the Riemann zeta function, while competitor OpenAI has deployed its Astra model to solve Erdos problems and advance theoretical computer science

5

.

Today's Top Stories

© 2026 TheOutpost.AI All rights reserved