AI firms turn to old books to escape AI-generated text contamination in training data

Reviewed byNidhi Govil

2 Sources

Share

AI companies are buying millions of pre-2022 printed books as clean training data to avoid the machine-generated text flooding the web. Data broker ISBNdb sources physical books in bulk, destroying them during scanning while keeping client identities secret under strict NDAs to avoid reputational damage.

AI Firms Seek Slop-Free Training Data in Pre-2022 Books

AI firms are turning to an unexpected source for clean training data: old books gathering dust on shelves

1

. Data broker ISBNdb, which claims to operate the world's largest book database, is now sourcing physical books in bulk specifically for AI companies to scan into AI training data

2

. The company's pitch is straightforward: books printed before 2022 represent "dense, edited, authoritative" high-quality curated knowledge that predates the large language model era

1

. This timing matters because it guarantees the text is free from AI-generated text contamination, a growing problem that threatens model training quality.

Source: 404 Media

Source: 404 Media

The urgency behind this AI data sourcing strategy stems from a documented phenomenon called model collapse. When AI models train on synthetic outputs generated by earlier models, each generation degrades slightly, producing progressively worse results

1

. The open web no longer offers reliable human-authored text, as machine-generated content now constitutes a fast-growing share of online material. Pre-2022 print runs, by contrast, provide a fixed record that nobody can quietly rewrite after publication.

The Data Poisoning Challenge Accelerates Book Acquisitions

Beyond avoiding their own AI slop, labs face another threat: deliberate data poisoning by authors fighting back against unauthorized use of their work

1

. Tools like Nightshade allow creators to embed characters that appear normal to human readers but corrupt model training. Research from Anthropic suggests that as few as 250 to 500 specially crafted documents can plant a backdoor in a corpus containing trillions of tokens

1

. Old books written before these defensive tools existed sidestep this problem entirely, offering AI companies a provenance advantage with a clean chain of custody for legal teams.

Destroying Millions of Books Under Strict Secrecy

The process of converting old books into slop-free training data comes with a destructive reality that ISBNdb acknowledges openly to potential clients. Scanning at scale typically requires workers to slice off book spines so loose pages can feed through machines, a method chosen because it's faster and cheaper than careful alternatives

1

. To address what the company calls the "optics problem," ISBNdb offers buyers strict NDAs on every engagement, promising that client names are never disclosed. The company's own marketing materials acknowledge the reputational damage at stake: "'AI company destroys two million books' is not a headline that generates sympathy"

1

. One suggested workaround involves reframing the destruction as digital preservation.

Copyright Battles and Legal Precedents Shape the Market

This book-buying strategy intersects directly with ongoing copyright battles in the AI industry. A US judge recently approved Anthropic's $1.5 billion settlement over pirated books and ruled that training on purchased, scanned books qualifies as fair use

1

. Part of the legal reasoning held that because the scanning process destroyed each print original, one legal copy simply replaced another. This court decision effectively validates buying and shredding paper as the legally blessed route for AI companies. Ingram, the largest book distributor in the US, has already warned publishers about these bulk purchases and offered them a way to opt out

1

. The development reveals an irony: the industry that promised to digitize human knowledge now pays to buy it up and pulp it, one truckload at a time. As AI models grow more sophisticated, the scarcest resource may not be computing power or algorithms, but rather sentences that no machine has ever touched.

Today's Top Stories

© 2026 TheOutpost.AI All rights reserved