AI Companies Buy, Scan, and Destroy Millions of Books to Train Large Language Models

Reviewed byNidhi Govil

5 Sources

Share

AI companies are purchasing millions of physical books and destroying them through destructive scanning to obtain high-quality human-written training data. The practice targets rare and out-of-print titles, raising ethical concerns about cultural heritage loss while creating a lucrative market for booksellers.

News article

AI Companies Turn to Physical Books for Clean Training Data

AI companies are aggressively purchasing millions of physical books to train large language models, employing destructive scanning methods that permanently destroy the originals. This practice emerged as developers seek high-quality human-written training data untainted by AI-generated content flooding the internet

1

. The process involves removing book spines, feeding pages through industrial scanners, and recycling or discarding what remains

3

.

The demand stems from a phenomenon researchers call model collapse—when AI models trained on content produced by earlier models begin reinforcing mistakes while losing variety and less common information from original human data

1

. Physical books printed before the current AI boom offer a dependable record of human writing, providing dense, edited, authoritative content that web crawls cannot replicate.

Anthropic Project Panama Reveals Industry Practices

Newly unsealed court filings exposed Anthropic's secretive "Project Panama," an internal initiative described in company documents as an effort to "destructively scan all the books in the world"

4

. The company spent tens of millions of dollars acquiring millions of used books, using hydraulic cutting machines to remove bindings before scanning pages with industrial imaging equipment to train its Claude models

4

.

Internal planning documents stated: "We don't want it to be known that we are working on this." Anthropic hired Tom Turvey, a former Google executive who helped create Google Books, to lead the effort

4

. One Anthropic co-founder wrote internally that books could teach AI models "how to write well" instead of imitating "low quality internet speak"

4

. The company initially explored sourcing from libraries and used bookstores before purchasing large batches from retailers like Better World Books. Proposals indicated plans to digitize between 500,000 and 2 million books in six months

4

.

ISBNdb and the Hidden Supply Chain

ISBNdb, a company operating an online book database, briefly advertised services to source up to 1 million books per order for AI developers before removing the pages following media attention

1

. The company offered books "tailored to your LLM training needs," including older, specialized, rare, and out-of-print titles gathered from used bookstores and catalogs

1

. Marketing materials promised a "strict NDA on every engagement," ensuring clients' identities, strategies, and acquisition targets would not be disclosed

1

.

ISBNdb argued that "print books from the pre-LLM era are structurally guaranteed to be free of this contamination," referring to AI-generated content

5

. The company acknowledged reputational risks, stating on its website: "The optics problem is real. 'AI company destroys two million books' is not a headline that generates sympathy"

5

. ISBNdb later told media outlets the service was exploratory and never launched, though the archived pages reveal detailed offerings for bulk book purchases for AI

2

.

Bulk Orders Alarm Independent Booksellers

Booksellers across the United States and Europe report receiving unusual bulk orders for obscure and out-of-print titles. Dutch antiquarian bookseller Pieter de Vries received an email from someone identifying as "Natalia" with company "2077AI," requesting quotes for more than 3,000 ISBNs, mostly published between 2020 and 2021 by academic publishers including Emerald Publishing, Elsevier, Wiley, Routledge and Oxford University Press

2

. De Vries initially dismissed it as spam or phishing, and several other Dutch antiquarian booksellers received identical requests

2

.

One American bookseller told media that weekly sales jumped from fewer than 20 books to hundreds in April, with buyers selecting seemingly random titles that all carried ISBN numbers

5

. The bookseller expressed mixed feelings: "It benefits me financially as well as by clearing out old inventory that is otherwise unlikely to sell. On the other hand, I don't like the end-use, and I don't like that uncommon books are being pulped"

5

. Similar buying patterns emerged in Germany and Switzerland, where local media reported secondhand booksellers received bulk orders for highly specialized books inconsistent with traditional collecting

2

.

Copyright Battles and Fair Use Rulings

The digitization of rare books for AI training has triggered copyright lawsuits against multiple companies. In June, a US federal judge ruled that using books to train AI models could qualify as fair use under the first-sale doctrine because the process is "transformative"—converting books into digital form rather than redistributing them as new copies

4

. However, the judge found Anthropic could still face liability over how it acquired books by downloading pirated copies from shadow libraries like LibGen and Pirate Library Mirror

4

.

Anthropically later agreed to pay $1.5 billion to settle claims related to book acquisition methods, while maintaining the settlement concerned acquisition rather than the legality of AI training itself

4

. Court filings revealed that before launching Anthropic Project Panama, employees downloaded books from shadow libraries. Anthropic stated it never trained a commercial AI model using the LibGen dataset and did not use Pirate Library Mirror to train any complete AI model

4

.

Separate lawsuits involving Meta, OpenAI and Google highlight how leading AI companies sought access to millions of books. Meta employees discussed using LibGen with internal messages showing copyright concerns, with one engineer writing: "Torrenting from a corporate laptop doesn't feel right"

4

. OpenAI acknowledged downloading LibGen but told a court it deleted files before releasing ChatGPT

4

. Most copyright cases involving AI companies, authors and publishers remain unresolved in US courts, leaving broader legal boundaries for training AI on copyrighted material uncertain

4

.

Historical Context: From Google Books to AI Training

Mass book scanning predates current AI applications. Project Gutenberg began creating electronic versions of public-domain works in 1971 and now offers more than 75,000 free ebooks

1

. Google Books launched in 2004, scanning books in partnership with libraries and publishers to create searchable online content. By 2019, Google assembled a collection exceeding 40 million books in over 400 languages

1

. Those scans helped form HathiTrust, a digital repository created by research libraries in 2008

1

.

Authors sued Google for scanning copyrighted books without permission, but a federal appeals court ruled in 2015 that the project qualified as fair use

1

. The court determined that creating a searchable index and displaying limited snippets gave books a new purpose without providing readers with a replacement for originals. However, AI companies' bulk book purchases for AI differ from earlier digitization efforts by preservation-focused organizations like the Internet Archive, which has digitized more than 25 million books since 2006 using nondestructive scanning

1

.

What This Means for Cultural Heritage and Future AI Development

The practice raises ethical concerns about permanently destroying scarce literary works and cultural heritage for AI training data. Rare and out-of-print titles being pulped may represent the last remaining physical copies of certain works

5

. The use of confidentiality agreements and intermediaries suggests AI companies recognize the reputational risks while continuing to pursue high-quality human-written training data through bulk book purchases for AI

1

.

As the internet increasingly fills with AI-generated content, demand for pre-AI era physical books will likely intensify. Watch for additional legal challenges testing the boundaries of fair use and copyright in AI contexts, potential regulations governing the digitization of rare books, and whether preservation efforts can keep pace with destructive scanning practices. The tension between technological advancement and cultural preservation will shape how AI companies source training data going forward, with implications for both the AI industry and the literary world.

Today's Top Stories

© 2026 TheOutpost.AI All rights reserved