4 Sources
[1]
Anthropic destroyed millions of print books to build its AI models
On Monday, court documents revealed that AI company Anthropic spent millions of dollars physically scanning print books to build Claude, an AI assistant similar to ChatGPT. In the process, the company cut millions of print books from their bindings, scanned them into digital files, and threw away
[2]
Anthropic destroyed millions of physical books to train its AI, court documents reveal
Serving tech enthusiasts for over 25 years. TechSpot means tech analysis and advice you can trust. WTF?! Generative AI has already faced sharp criticism for its well-known issues with reliability, its massive energy consumption, and the unauthorized use of copyrighted material. Now, a recent
[3]
Anthropic Shredded Millions of Physical Books to Train its AI
Today in schnozz-smashing on-the-nose metaphors for the AI industry's rapacious destruction of the arts: exactly how Anthropic gathered the data it needed to train its Claude AI model. As Ars Technica reports, the Google-backed startup didn't just crib from millions of copyrighted books, a
[4]
Anthropic trashed millions of books to train its AI
Anthropic physically scanned millions of print books to train its AI assistant, Claude, subsequently discarding the originals, as revealed in court documents, according to Ars Tecnica. This extensive operation, detailed in a legal decision, involved the acquisition and destructive digitization of
Share
Copy Link
Anthropic, an AI company, destroyed millions of physical books to train its AI model Claude, sparking debates on data acquisition methods, copyright, and ethics in AI development.
In a shocking revelation, court documents have exposed that AI company Anthropic engaged in the destruction of millions of physical books to train its AI model, Claude. This controversial practice, aimed at acquiring high-quality training data, has ignited debates on the ethics and legality of AI development methods
1
.
Source: Ars Technica
Anthropic's approach involved purchasing millions of physical books, cutting them from their bindings, scanning them into digital files, and discarding the originals. This process, known as destructive scanning, was implemented on an unprecedented scale. The company hired Tom Turvey, former head of partnerships for Google Books, to spearhead this operation in February 2024
1
.U.S. District Judge William Alsup ruled that Anthropic's destructive scanning operation qualified as fair use. This decision was based on several factors:
The judge compared the process to "conserving space" through format conversion and deemed it transformative
2
.The case highlights the AI industry's insatiable appetite for high-quality text data. Large language models (LLMs) like Claude require billions of words for training, with the quality of input directly impacting the model's capabilities
1
.
Source: Futurism
Anthropic's approach exploited the first-sale doctrine, which allows buyers to do what they want with their purchases without copyright holder intervention. This legal workaround enabled the company to avoid complex licensing negotiations with publishers
3
.The destruction of millions of books has raised ethical concerns within the archival and literary communities. Alternative methods for mass book digitization exist, such as those pioneered by the Internet Archive, which preserve physical volumes while creating digital copies
1
.Related Stories
Anthropic's partial legal victory allows it to train AI models on copyrighted books without notifying original publishers or authors. This ruling could have far-reaching consequences for the AI industry, potentially removing a significant hurdle in AI development
2
.Despite this ruling, Anthropic still faces a copyright trial in December for its earlier use of pirated ebooks. The company could be ordered to pay up to $150,000 per pirated work
2
.As the AI industry grapples with data scarcity and copyright issues, companies are exploring various approaches. OpenAI and Microsoft recently announced a collaboration with Harvard's libraries to train AI models on nearly 1 million public domain books, demonstrating a more ethically sound approach to data acquisition
4
.Summarized by
Navi
[4]
23 Jul 2026•Policy and Regulation
30 Jul 2026•Policy and Regulation

12 Aug 2026•Technology
