7 Sources
[1]
AI firms are quietly buying and destroying millions of printed books to train their models
Serving tech enthusiasts for over 25 years. TechSpot means tech analysis and advice you can trust. What we know so far: A market built for libraries and booksellers is now feeding AI training, as companies seek large volumes of printed books that predate the generative AI boom. ISBNdb, which says it maintains the world's largest book database, now helps AI companies source physical books in bulk. The company says it can coordinate purchases ranging from 1,000 to 1 million books, drawing from secondary markets where libraries, retailers, and individuals sell used inventory. Its pitch is simple: older books offer cleaner data. "The world's best AI training data is sitting on a shelf," the company says on its website. "Books represent curated, peer-reviewed, domain-specific human knowledge, structured in a way no web crawl can replicate. Dense, edited, authoritative." The focus is largely on books published before 2022, before synthetic text became widespread online. That matters because AI models trained on AI-generated content can degrade over time, a problem often referred to as model collapse. ISBNdb frames print as a way around that risk. "Print books from the pre-LLM era are structurally guaranteed to be free of this contamination," the company says. "Physical books published before this date [pre-2022] are structurally clean of modern poisoning tools." To make that usable at scale, ISBNdb combines its catalog with sourcing services that help buyers avoid duplicate titles and assemble large, varied datasets. It also emphasizes confidentiality. The site says every engagement is covered by a strict nondisclosure agreement and that the buyer's identity, strategy, and acquisition targets are kept confidential. That lack of transparency carries over into the marketplaces supplying the books. Sellers on platforms like Alibris and Biblio say they have seen a sharp increase in bulk orders in recent months, though the buyers are not identified. One bookseller who specializes in foreign-language titles said sales changed dramatically starting in April. "I personally have mixed feelings about all of this," the seller told 404 Media. "It benefits me financially as well as by clearing out old inventory that is otherwise unlikely to sell. I've been well-suited for these sales with inventory from overseas and foreign language books. On the other hand, I don't like the end-use, and I don't like that uncommon books are being pulped." The orders themselves stand out. Instead of focusing on a subject area, they appear scattered across topics and formats. The seller said the orders were unusual not just because of their size, but because they seemed random and ignored normal pricing patterns. The seller also suggested that the pattern looked like AI-related purchasing because of the scale of spending. Other sellers are seeing similar patterns. On an Alibris forum, one user asked, "Is it just me, or has there been an uptick in the number of AutoBuy orders since the tail end of last year?" Mike Feldman, director of client services at Alibris, replied, "We have a couple of new bulk buyers that are scooping up trade books so lots of sellers are getting lots of orders." There is no direct confirmation that these purchases are tied to specific AI companies, but recent court cases show how printed books are being used. In a lawsuit involving Anthropic, internal documents described a plan to buy and scan millions of books. In many cases, the books were destroyed during the process. That approach relies on high-speed destructive scanning, where the spine is cut and pages are fed into machines. It is faster and cheaper than scanning intact books. One vendor mentioned in the case, Datamation Information Services, offers both destructive and non-destructive options. A federal judge later ruled that Anthropic's process qualified as fair use. "Here, every purchased print copy was copied in order to save storage space and to enable searchability as a digital copy," Judge William Alsup wrote. "The print original was destroyed. One replaced the other. And, there is no evidence that the new, digital copy was shown, shared, or sold outside the company." He called the use "clearly transformative." ISBNdb points to that reasoning in its own materials. The company says buying paper books in bulk from the secondary market does not deprive creators of income, because the books have already been sold once. It also acknowledges the optics. "'AI company destroys two million books' is not a headline that generates sympathy." In another section, it adds, "Responsible physical sourcing is not book burning. It is the completion of a book's lifecycle: from tree to knowledge to tree again." Neither ISBNdb nor Anthropic responded to requests for comment. For AI developers, print offers a cleaner, more controlled source of training data. For sellers, the surge in demand is moving inventory that might otherwise sit for years. What remains unclear is how many rare or low-circulation books could be lost through destructive scanning.
[2]
AI firms are buying old books for slop-free training data
A data broker is selling AI labs on a simple idea: books printed before the chatbots are the only text a machine has not already contaminated. Getting at it means cutting the spine off millions of them, quietly. The pitch from data broker ISBNdb is blunt: the world's best AI training data is sitting on a shelf. The catch is that getting at it means slicing the spine off millions of books, then never admitting which lab paid for the job. AI has a pollution problem of its own making. So much of the web is now machine-written that models risk feeding on their own exhaust. One company thinks the antidote is sitting on a shelf: old, printed books, published before the chatbots arrived. 404 Media reported that ISBNdb now sources physical books in bulk for AI labs to scan into training data. The firm calls itself the world's largest book database. "The world's best AI training data is sitting on a shelf," its site says. Books, it adds, are "dense, edited, authoritative." Why old books The value is in the date. ISBNdb argues that books printed before 2022 predate the large language model era, so they cannot contain AI-generated text. That matters because of model collapse. It names the documented decline that sets in when models train on the synthetic output of earlier models. Each generation ends up a little worse than the last. The open web offers no such guarantee. A fast-growing share of online text is now machine-made, part of the same slop flood the labs helped create. A pre-2022 print run, by contrast, is a fixed, human-authored record that nobody can quietly rewrite. The poisoning arms race There is a second reason labs want clean paper. Authors have started to fight back with data poisoning. They borrow tricks from tools like Nightshade, lacing text with characters that a person reads normally but a model cannot. ISBNdb's own blog cites Anthropic research suggesting that as few as 250 to 500 crafted documents can plant a backdoor in a corpus of trillions of tokens. Pre-2022 books, written before any of these tools existed, sidestep the problem. ISBNdb sells that as provenance: buy the paper, keep the receipts, and your legal team holds a clean chain of custody. The optics problem There is a catch the company states plainly. Scanning at scale usually destroys the book. Workers slice off the spine so the loose pages feed through a machine, which is faster and cheaper than the careful alternative. So ISBNdb offers its clients secrecy. "Strict NDA on every engagement," its site says. Buyers' names are "never disclosed." The reason sits in its own marketing copy. "The optics problem is real," the site reads. "'AI company destroys two million books' is not a headline that generates sympathy." One suggested workaround is pure spin: reword the deed as digitally preserving the books. A familiar fight All of this runs straight into a copyright battle already raging. A US judge approved Anthropic's $1.5bn settlement over pirated books, and ruled that training on purchased, scanned books counts as fair use. Part of his reasoning was that the copying destroyed each print original, so one legal copy simply replaced another. Buying and shredding paper, in other words, is the route the courts bless. So the industry that promised to digitise human knowledge now pays to buy it up and pulp it, one lorry-load at a time. Ingram, the largest book distributor in the US, has already warned publishers and offered them a way to opt out. The scarcest thing in AI, it turns out, is a sentence that no machine ever touched.
[3]
AI Companies Are Buying Tons of Old Books Because They're Free of AI Slop
ISBNdb, a company that sources printed books for AI companies to turn into training data, tells clients "the optics problem is real." As AI companies search for more training data to improve their models, one company is offering old, printed books as an ideal source because they are guaranteed to be free of the very AI slop AI companies are producing. "The world's best AI training data is sitting on a shelf," ISBNdb, a company that produces what it claims is "the world's largest book database," and that offers high-volume book acquisition services for AI companies, says on its site. "Books represent curated, peer-reviewed, domain-specific human knowledge, structured in a way no web crawl can replicate. Dense, edited, authoritative."
[4]
AI Companies Are Buying Antique Books, Ingesting Their Contents to Train Models, and Then Destroying Them at Incredible Scale
Can't-miss innovations from the bleeding edge of science and tech Traditionally, books have been good for two things: reading, and looking nice on a shelf. But AI companies are interested in neither. To those building large language models, books are nothing more than fodder to be devoured en masse before being spit out like fishbone. Often, they're happy to use digital books -- or even better, pirated digital books, as Meta has been accused of doing, and as Anthropic was forced to pay a $1.5 billion settlement to authors for also doing. But many companies, including Anthropic, have turned to ingesting physical books instead, which they can buy countless used copies of on the cheap. According to the settled lawsuit, Anthropic used a hydraulic powered cutting machine to neatly remove the pages from the books it procured from book resellers and then scanned them using industrial-grade imaging equipment. In other words, it was literally ripping off authors' books to train its AI. This process took advantage of a legal concept known as first-sale doctrine, which allows a buyer to do what they want with a purchase without the original copyright holder's say-so. And since Anthropic was turning the original physical texts into digital ones -- rather than redistributing them as new copies -- a judge found this to be "transformative," and therefore protected by fair use. Now, as 404 Media reports, this practice has become prevalent enough that even well-established book sellers are looking to cash in on the AI boom. One called ISBNdb, which boasts the "world's largest book database," extolls that the "world's best AI training data is setting on a shelf," upholding these physical texts as uncorrupted by shoddy AI writing that's already polluted so much of the internet (and indeed, newer books). "Print books from the pre-LLM era are structurally guaranteed to be free of this contamination. That alone is a significant advantage," it explains in an article on its website, as quoted by 404. "Physical books published before this date [pre-2022] are structurally clean of modern poisoning tools." Once focused on helping libraries, distributors, and book shops find and sell books, ISBNdb now helps AI companies bulk-buy anywhere between 1,000 to one million books per order, according to 404. As an added bonus, it also promises AI companies that it'll keep their purchases under wraps -- nobody wants to end up in the spotlight like Anthropic and Meta, obviously -- while clearly sounding aware about how incredibly shady the practice sounds. "The optics problem is real," ISBNdb's site says. "'AI company destroys two million books' is not a headline that generates sympathy." One small book seller said that in April, he suddenly went from selling no more than 20 books a week to hundreds, and he's almost certain that the customers are AI labs, noting the random selection of the books and how they all have ISBNs. He added that his inventory is full with rare and out of print books, meaning that an AI company could be destroying some of the few remaining copies that can be found. "I personally have mixed feelings about all of this," the bookseller told 404. "It benefits me financially as well as by clearing out old inventory that is otherwise unlikely to sell. I've been well-suited for these sales with inventory from overseas and foreign language books. On the other hand, I don't like the end-use, and I don't like that uncommon books are being pulped." More antique books may be endangered. 404 noted how rare booksellers in the Netherlands are being inundated with bulk purchases they believe are being made by AI companies, though they can't know for sure. Similar suspicions are being felt all throughout the industry. Services like ISBNdb facilitate these bulk orders, keeping the AI buyers anonymous as promised. There may be strong hints thats a bulk order is for an AI lab, but booksellers can, for the most part, merely speculate.
[5]
AI Book Burning? Companies Are Destroying Millions of Books to Feed Chatbots
A federal judge ruled that destructive scanning of legally purchased books can qualify as fair use, even as separate copyright litigation continues. Like a scene out of the classic dystopian novel "Fahrenheit 451," some AI companies aren't just reading books -- they're destroying them. As developers race to build more powerful AI models, and lawsuits over copyright mount, a cottage industry has sprung up to supply them with millions of physical books that are stripped apart, scanned into training datasets, and discarded. First reported by 404 Media, AI companies are using intermediaries to acquire books at industrial scale anonymously. Companies specializing in bulk sourcing advertise their ability to locate hundreds of thousands of titles while promising confidentiality for AI clients, reflecting the sensitivity surrounding the practice. Critics say the buying spree is driven by AI companies racing to preserve human-authored knowledge before it is diluted by AI-generated text, often called "AI slop." Books published before the rise of generative AI in 2023 are especially valuable because they provide high-quality training data written entirely by humans. The surge in demand is already reshaping the used-book market. One unnamed bookseller told 404 Media that weekly sales climbed from roughly 20 books to several hundred after AI buyers entered the market. While the increase has been profitable, he said he worries uncommon and out-of-print books are being permanently lost after they are scanned and destroyed. "It benefits me financially as well as by clearing out old inventory that is otherwise unlikely to sell," the bookseller told 404 Media. "I've been well suited for these sales with inventory from overseas and foreign language books. On the other hand, I don't like the end-use, and I don't like that uncommon books are being pulped." The practice mirrors Anthropic's "Project Panama," which digitized millions of books through destructive scanning. Last summer, in the copyright lawsuit Bartz v. Anthropic PBC, a federal judge in San Francisco ruled that scanning legally purchased physical books into digital copies, even when the originals were destroyed, constituted transformative fair use. Federal judges later issued similar fair use rulings in separate copyright cases involving OpenAI and Meta. However, in a separate case, a federal judge in the same district this week approved a $1.5 billion copyright settlement requiring Anthropic to pay thousands of authors about $3,000 per book after the company used pirated copies of their works to train Claude. In response to the growing backlash, AI developers, including Elon Musk, have spoken out against the practice and said companies should maintain the books being scanned. "I've asked the SpaceXAI team to preserve any rare books in a library and scan them the hard way vs just cutting off the spine and scanning," Musk wrote on X.
[6]
AI companies are anonymously buying and destroying millions of books through middleman services to avoid headlines about AI companies buying and destroying millions of books
Booksellers say they're facing a sudden surge of orders for seemingly random rare books -- and fear they might be profiting from their destruction. AI isn't doing great on the PR front. Whether it's the exponentially increasing electricity consumption amidst worsening climate change, the data centers leaking toxic pollutants into neighboring communities, or the years-long parade of business leaders insisting the technology will replace humans in "most, if not all" professional fields, the AI industry has an excellent record of inspiring resentment and contempt. And, it seems, the AI companies know it. New reporting from 404 Media suggests that, in an effort to shield themselves from the PR fallout of their data harvesting, major AI providers have turned to third-party middleman services to anonymously source the physical books they've been stripping for training data at an industrial scale -- and destroying in the process. In an information ecosystem increasingly poisoned by regurgitated and recirculated AI slop, untainted samples of quality written text have, ironically, only become more valuable for the companies making them so hard to find. The AI model arms race requires an ever-widening set of training data, making physical books that predate the proliferation of LLM-generated text a precious source of pristine, human-authored material to imitate -- particularly if it's a book rare enough to be absent from your competitor's datasets. Unfortunately, a book is only valuable for an AI company until its last page has been scanned, as demonstrated by a court ruling in 2025 revealing that Anthropic had carved up, de-spined, scanned, and ultimately discarded millions of print books in the process of training its Claude models. Destroying books, it turns out, isn't just cheaper than maintaining them: The presiding judge also ruled that it's transformative enough to constitute fair use under Section 107 of the Copyright Act. But while a judge might have declared it legally defensible, The Washington Post reported that internal documents -- unsealed during the author-initiated copyright lawsuit that ended in the company paying a $1.5 billion settlement -- indicated Anthropic was well aware that 'arguably legal' and 'cool to do' are very different things. "We don't want it to be known that we are working on this," Anthropic said in internal planning documents, understanding that mulching millions of books just so its chatbot didn't sound like "low quality internet speak" might not have been a popular move. That's why, as 404 Media now reports, Anthropic and other AI providers are hiring middleman companies to buy those books instead -- companies like ISBNdb, which until earlier today featured a now-removed page on its site offering printed book sourcing services, "tailored to your LLM training needs, delivered at the scale AI demands." In an archived version of the page, ISBNdb boasted about its ability to secure books "scattered across library shelves, used bookstores, and out-of-print catalogs," with companies able to purchase up to 1 million books per order. One of the website's blog posts about the service -- also updated today -- featured a now-removed promise of a "strict NDA on every engagement," ensuring that "your identity, strategy, and acquisition targets are never disclosed." After all, the blog post still reads, "'AI company destroys two million books' is not a headline that generates sympathy." In a news update, ISBNdb says it removed the page describing the AI book-sourcing service because it was simply "part of exploring demand, and we've chosen to pivot away from that direction." ISBNdb might be pivoting, but there's clearly plenty of that demand to go around. 404 quotes a bookseller specializing in rare and low-circulation books who said he and other book dealers have seen an unprecedented surge in sales since April. He's now fulfilling orders for hundreds of books a week when he'd previously hoped to clear 20. There's no commonality between the seemingly random books being bulk-ordered, the bookseller said, except the fact that they all have ISBNs, potentially indicating that the books had been identified by the purchaser from an ISBN-based database. Netherlands book dealers experiencing a similar boom in rare book orders noticed the same trend earlier this year. "I personally have mixed feelings about all of this. It benefits me financially as well as by clearing out old inventory that is otherwise unlikely to sell. I've been well-suited for these sales with inventory from overseas and foreign language books," the bookseller said. "On the other hand, I don't like the end-use, and I don't like that uncommon books are being pulped."
[7]
AI companies are secretly destroying rare, irreplaceable books
AI companies are buying up rare and out-of-print books by the truckload, cutting off the bindings, scanning every page, and then pulping what's left. Some of these books have almost no surviving copies anywhere else. If you've ever sold an old book, inherited a box of them, or just like the idea that a rare title survives somewhere, this should bother you. Once one of these books gets fed into the machine, it's gone. There's no digital backup sitting in a library. There's no second copy waiting in someone's attic. It's simply destroyed. How this is actually happening According to reporting from 404 Media and others, the process is called destructive scanning. High-speed machines slice the spine off a book, feed the loose pages through a scanner, and then the paper gets thrown out or pulped. Companies want books published before 2022, since those are less likely to contain AI-generated text that could pollute the data they're collecting. Booksellers in the US, Canada, and Europe say they're getting bulk orders that don't make sense for normal customers. One seller told reporters he used to sell about 20 books a week. Since April, he's regularly selling hundreds. The orders are random and tied to ISBNs, not to which books are rare or hard to find. Some of this is happening under strict secrecy. A service called ISBNdb reportedly offers AI companies non-disclosure agreements so their identities never get attached to the purchases. Its own website even acknowledges the problem, saying a headline like "AI company destroys two million books" doesn't exactly earn public sympathy. A federal judge has already ruled that scanning purchased books and destroying the originals afterward counts as fair use under copyright law. That ruling is part of why this practice has been able to scale up so quickly. Reports also point to Anthropic having run a large-scale version of this earlier, under an internal project reportedly called Project Panama. What this means if you've got old books lying around Nobody checks whether a book is one of the last copies on Earth before it goes through the scanner. If you or a relative has a box of old or unusual books gathering dust, it might be worth getting them appraised or checked out by a local library or rare book dealer before selling them off cheap. A book that looks worthless to you could be one of the only copies left, and once it's sold into this pipeline, there's no getting it back.
Share
Copy Link
AI firms are quietly bulk-purchasing millions of physical books published before 2022 to train their models, often destroying them through destructive scanning. ISBNdb now coordinates purchases ranging from 1,000 to 1 million books, marketing pre-AI era texts as structurally guaranteed to be free of synthetic content contamination. While a federal judge ruled this practice qualifies as fair use, the surge is reshaping the used-book market and raising ethical concerns about rare titles being permanently lost.
A market originally built for libraries and booksellers is now feeding the appetite of AI companies desperate for clean training data. ISBNdb, which maintains what it claims is the world's largest book database, has pivoted to help AI firms source physical books in bulk—coordinating purchases ranging from 1,000 to 1 million books at a time
1
. The company's pitch is straightforward: "The world's best AI training data is sitting on a shelf," emphasizing that printed books represent "curated, peer-reviewed, domain-specific human knowledge, structured in a way no web crawl can replicate"3
.Source: TechSpot
The focus centers overwhelmingly on books published before 2022, before AI-generated text contamination became widespread online. This temporal boundary matters because AI models trained on synthetic content can suffer from model collapse—a documented degradation where each generation performs worse than the last
2
. ISBNdb frames print as a solution to this risk, stating that "print books from the pre-LLM era are structurally guaranteed to be free of this contamination" and are "structurally clean of modern poisoning tools" [1](https://www.techspot.com/news/113277-ai-firms-qui etly-buying-destroying-millions-printed-books.html).To make these books usable at scale for large language models, AI companies employ destructive scanning—a process where the spine is cut and pages are fed through high-speed machines. This method is faster and cheaper than scanning intact books, but it permanently destroys the originals
1
. Internal documents from Anthropic's Project Panama described plans to buy and scan millions of books, with many destroyed during the process4
. According to the settled lawsuit, Anthropic used "a hydraulic powered cutting machine to neatly remove the pages from the books it procured from book resellers and then scanned them using industrial-grade imaging equipment"4
.
Source: Futurism
The surge in bulk-purchasing books has dramatically altered the used-book market. One bookseller specializing in foreign-language titles reported that sales changed dramatically starting in April, jumping from roughly 20 books per week to several hundred
4
5
. "I personally have mixed feelings about all of this," the seller told 404 Media. "It benefits me financially as well as by clearing out old inventory that is otherwise unlikely to sell. On the other hand, I don't like the end-use, and I don't like that uncommon books are being pulped"1
.ISBNdb emphasizes strict confidentiality in its data acquisition strategies, stating that every engagement is covered by a nondisclosure agreement and that "the buyer's identity, strategy, and acquisition targets are kept confidential"
1
. This secrecy reflects what the company itself acknowledges as an optics problem. "'AI company destroys two million books' is not a headline that generates sympathy," ISBNdb's website plainly states1
4
.Sellers on platforms like Alibris and Biblio report seeing sharp increases in bulk orders in recent months, though buyers remain unidentified. On an Alibris forum, one user asked about an uptick in AutoBuy orders since late last year. Mike Feldman, director of client services at Alibris, confirmed: "We have a couple of new bulk buyers that are scooping up trade books so lots of sellers are getting lots of orders"
1
. The orders stand out not just for their size but because they appear scattered across topics and formats, ignoring normal pricing patterns.The practice has sparked significant copyright litigation. Federal Judge William Alsup ruled that Anthropic's destructive scanning process qualified as fair use under the first-sale doctrine. "Here, every purchased print copy was copied in order to save storage space and to enable searchability as a digital copy," Judge Alsup wrote. "The print original was destroyed. One replaced the other. And, there is no evidence that the new, digital copy was shown, shared, or sold outside the company." He called the use "clearly transformative"
1
4
.However, in a separate case, a federal judge approved a $1.5 billion settlement requiring Anthropic to pay thousands of authors approximately $3,000 per book after the company used pirated copies to train Claude
5
. ISBNdb points to the fair use reasoning in its own materials, arguing that bulk-purchasing books from the secondary market doesn't deprive creators of income because the books have already been sold once1
.Related Stories
Beyond avoiding AI-generated text contamination, AI companies face another threat: data poisoning. Authors have started fighting back using tools like Nightshade, which lace text with characters that humans read normally but models cannot process correctly
2
. ISBNdb's blog cites Anthropic research suggesting that as few as 250 to 500 crafted documents can plant a backdoor in a corpus of trillions of tokens2
. Pre-2022 printed books, written before any of these tools existed, sidestep this problem entirely, offering what ISBNdb markets as provenance: buy the paper, keep the receipts, and legal teams hold a clean chain of custody.The practice has raised ethical concerns across the industry, particularly regarding out-of-print books and rare titles. Booksellers report that their inventory includes rare and out-of-print books, meaning AI companies could be destroying some of the few remaining copies
4
. Rare booksellers in the Netherlands are reportedly being inundated with bulk purchases believed to be from AI companies4
.Elon Musk has spoken out against the practice, stating: "I've asked the SpaceXAI team to preserve any rare books in a library and scan them the hard way vs just cutting off the spine and scanning"
5
. ISBNdb attempts to reframe the practice, stating: "Responsible physical sourcing is not book burning. It is the completion of a book's lifecycle: from tree to knowledge to tree again"1
. Ingram, the largest book distributor in the US, has already warned publishers and offered them a way to opt out2
.Summarized by
Navi
[1]
[2]
Yesterday•Entertainment and Society

30 Jul 2026•Policy and Regulation

26 Jun 2025•Technology

1
Science and Research

2
Technology

3
Technology
