2 Sources
[1]
AI firms are buying old books for slop-free training data
A data broker is selling AI labs on a simple idea: books printed before the chatbots are the only text a machine has not already contaminated. Getting at it means cutting the spine off millions of them, quietly. The pitch from data broker ISBNdb is blunt: the world's best AI training data is sitting on a shelf. The catch is that getting at it means slicing the spine off millions of books, then never admitting which lab paid for the job. AI has a pollution problem of its own making. So much of the web is now machine-written that models risk feeding on their own exhaust. One company thinks the antidote is sitting on a shelf: old, printed books, published before the chatbots arrived. 404 Media reported that ISBNdb now sources physical books in bulk for AI labs to scan into training data. The firm calls itself the world's largest book database. "The world's best AI training data is sitting on a shelf," its site says. Books, it adds, are "dense, edited, authoritative." Why old books The value is in the date. ISBNdb argues that books printed before 2022 predate the large language model era, so they cannot contain AI-generated text. That matters because of model collapse. It names the documented decline that sets in when models train on the synthetic output of earlier models. Each generation ends up a little worse than the last. The open web offers no such guarantee. A fast-growing share of online text is now machine-made, part of the same slop flood the labs helped create. A pre-2022 print run, by contrast, is a fixed, human-authored record that nobody can quietly rewrite. The poisoning arms race There is a second reason labs want clean paper. Authors have started to fight back with data poisoning. They borrow tricks from tools like Nightshade, lacing text with characters that a person reads normally but a model cannot. ISBNdb's own blog cites Anthropic research suggesting that as few as 250 to 500 crafted documents can plant a backdoor in a corpus of trillions of tokens. Pre-2022 books, written before any of these tools existed, sidestep the problem. ISBNdb sells that as provenance: buy the paper, keep the receipts, and your legal team holds a clean chain of custody. The optics problem There is a catch the company states plainly. Scanning at scale usually destroys the book. Workers slice off the spine so the loose pages feed through a machine, which is faster and cheaper than the careful alternative. So ISBNdb offers its clients secrecy. "Strict NDA on every engagement," its site says. Buyers' names are "never disclosed." The reason sits in its own marketing copy. "The optics problem is real," the site reads. "'AI company destroys two million books' is not a headline that generates sympathy." One suggested workaround is pure spin: reword the deed as digitally preserving the books. A familiar fight All of this runs straight into a copyright battle already raging. A US judge approved Anthropic's $1.5bn settlement over pirated books, and ruled that training on purchased, scanned books counts as fair use. Part of his reasoning was that the copying destroyed each print original, so one legal copy simply replaced another. Buying and shredding paper, in other words, is the route the courts bless. So the industry that promised to digitise human knowledge now pays to buy it up and pulp it, one lorry-load at a time. Ingram, the largest book distributor in the US, has already warned publishers and offered them a way to opt out. The scarcest thing in AI, it turns out, is a sentence that no machine ever touched.
[2]
AI Companies Are Buying Tons of Old Books Because They're Free of AI Slop
ISBNdb, a company that sources printed books for AI companies to turn into training data, tells clients "the optics problem is real." As AI companies search for more training data to improve their models, one company is offering old, printed books as an ideal source because they are guaranteed to be free of the very AI slop AI companies are producing. "The world's best AI training data is sitting on a shelf," ISBNdb, a company that produces what it claims is "the world's largest book database," and that offers high-volume book acquisition services for AI companies, says on its site. "Books represent curated, peer-reviewed, domain-specific human knowledge, structured in a way no web crawl can replicate. Dense, edited, authoritative."
Share
Copy Link
AI companies are buying millions of pre-2022 printed books as clean training data to avoid the machine-generated text flooding the web. Data broker ISBNdb sources physical books in bulk, destroying them during scanning while keeping client identities secret under strict NDAs to avoid reputational damage.
AI firms are turning to an unexpected source for clean training data: old books gathering dust on shelves
1
. Data broker ISBNdb, which claims to operate the world's largest book database, is now sourcing physical books in bulk specifically for AI companies to scan into AI training data2
. The company's pitch is straightforward: books printed before 2022 represent "dense, edited, authoritative" high-quality curated knowledge that predates the large language model era1
. This timing matters because it guarantees the text is free from AI-generated text contamination, a growing problem that threatens model training quality.
Source: 404 Media
The urgency behind this AI data sourcing strategy stems from a documented phenomenon called model collapse. When AI models train on synthetic outputs generated by earlier models, each generation degrades slightly, producing progressively worse results
1
. The open web no longer offers reliable human-authored text, as machine-generated content now constitutes a fast-growing share of online material. Pre-2022 print runs, by contrast, provide a fixed record that nobody can quietly rewrite after publication.Beyond avoiding their own AI slop, labs face another threat: deliberate data poisoning by authors fighting back against unauthorized use of their work
1
. Tools like Nightshade allow creators to embed characters that appear normal to human readers but corrupt model training. Research from Anthropic suggests that as few as 250 to 500 specially crafted documents can plant a backdoor in a corpus containing trillions of tokens1
. Old books written before these defensive tools existed sidestep this problem entirely, offering AI companies a provenance advantage with a clean chain of custody for legal teams.The process of converting old books into slop-free training data comes with a destructive reality that ISBNdb acknowledges openly to potential clients. Scanning at scale typically requires workers to slice off book spines so loose pages can feed through machines, a method chosen because it's faster and cheaper than careful alternatives
1
. To address what the company calls the "optics problem," ISBNdb offers buyers strict NDAs on every engagement, promising that client names are never disclosed. The company's own marketing materials acknowledge the reputational damage at stake: "'AI company destroys two million books' is not a headline that generates sympathy"1
. One suggested workaround involves reframing the destruction as digital preservation.Related Stories
This book-buying strategy intersects directly with ongoing copyright battles in the AI industry. A US judge recently approved Anthropic's $1.5 billion settlement over pirated books and ruled that training on purchased, scanned books qualifies as fair use
1
. Part of the legal reasoning held that because the scanning process destroyed each print original, one legal copy simply replaced another. This court decision effectively validates buying and shredding paper as the legally blessed route for AI companies. Ingram, the largest book distributor in the US, has already warned publishers about these bulk purchases and offered them a way to opt out1
. The development reveals an irony: the industry that promised to digitize human knowledge now pays to buy it up and pulp it, one truckload at a time. As AI models grow more sophisticated, the scarcest resource may not be computing power or algorithms, but rather sentences that no machine has ever touched.Summarized by
Navi
[1]
26 Jun 2025•Technology

13 Jun 2025•Technology

06 Feb 2025•Technology

1
Technology

2
Science and Research

3
Policy and Regulation
