16 Sources
[1]
Are AI Companies Really Destroying Books?
Despite the panic online, there are plenty of unanswered questions. Last week, the internet raged as social-media posts and news articles accused AI companies of "destroying the world's books" -- including "millions of rare" ones -- by chopping off their spines to scan them more easily, and discarding them afterward. One article said the news was "sparking concerns that the last remaining copies of out-of-print texts are being destroyed." The investor and AI skeptic Michael Burry called the practice of sacrificing rare books "evil incarnate." Even Elon Musk weighed in, posting that he had asked the engineers training xAI's models to "preserve any rare books in a library and scan them the hard way." The destruction of physical books in the process of training AI has been public knowledge for more than a year, but debate over the practice intensified after an article by 404 Media, which described a large number of orders coming to booksellers, suggested that rare books might be included, and named a book-database company called ISBNdb as the possible buyer. Lost in the outrage, however, are several unanswered questions. The 404 Media story said ISBNdb had advertised bulk buying of books for AI companies, but acknowledged that no evidence directly connects the company to the recent wave of purchases that booksellers have described. And although the story calls out rare-book sellers, it's unclear how rare most of the books being bought actually are, or whether they're of great value. Despite that, the conversation rapidly intensified into a panic. So what is actually going on? And who is actually responsible? Read: Is This What Comes After AI Slop? Many of the specifics of this situation are obscured by hard-to-trace purchasing histories and the silence of AI companies and other potential buyers, so let's lay out what we know. We know that at least one major AI company has used physical books to train their models. Documents unsealed earlier this year in a lawsuit against Anthropic reveal that in the spring of 2024, the AI giant bought millions of new and used books and destroyed them in order to scan the pages to train its AI models. Anthropic and other AI companies are likely keen on books published before the launch of ChatGPT, in 2022, which are all but guaranteed to be free of AI-generated text. Although AI-generated text is readable -- if often unpleasant -- to humans, it can be devastating to AI models: When trained on their own output, they can degrade significantly in quality in a phenomenon known as "model collapse." We also know that something mysterious is afoot. Someone -- if not more than one group or company -- has been snapping up massive volumes of books from shops around the world. "What's with the online purchasing bonanza of used books right now?" asked one bookseller on Reddit three months ago. Several people who identified themselves as booksellers responded with their own stories of unusually large orders. Charlie Becker, the manager of Becker's Books, in Houston, told me that 95 of the past 100 books ordered through one of his online-selling platforms went to a single buyer. One of the orders was for 70 books. (To avoid breaking the terms of service of his selling platform, Becker would not tell me the name of the buyer.) But contrary to many news reports, the old books bought from these sellers do not appear to be "rare" in the typical sense of the word. According to Becker and the online discussions I read, the books in these orders are almost entirely nonfiction, were mostly published from the 1970s to the 1990s, and were written in a variety of languages. Becker wrote in a blog post that he has received orders for "stuff like The Insider's Guide to Metro Denver from 1995 or How to Use Corel WordPerfect 1991." These belong to a category of books that you could call rare because a small number of copies remain, but they may also be out-of-date, undesirable to readers, and not valuable per se. In some cases, booksellers have noted, the cost of shipping exceeds the books' resale value. By most accounts, AI companies want a large volume of text for as little cost as possible. It would be strange if they were intentionally paying premium prices for first editions, unusual printings of cult classics, or yellowed books of esoteric knowledge. Yet it's possible: The pursuit of training data that their competitors don't have could lead them to such titles. We also have a good sense of who's behind the book buying -- and it seems unlikely to be ISBNdb.com, the subject of the 404 Media story. For one thing, no bookseller I could find in online discussions had mentioned an order from ISBNdb. The company maintains a public book database and charges for access to it, but does not appear to buy or sell any books. Things get confusing because, in March, the company added a page to its website that offered AI developers the ability to buy physical books, "up to 1 million titles per order," with a nondisclosure agreement to preserve anonymity. When I asked ISBNdb about the service, I received an email from a customer-support representative claiming that "ISBNdb has never purchased, scanned, or destroyed a book -- for AI training or anything else. We don't train AI models, and we never have. The page was up to explore demand for a service we never brought to life." In response to recent "concern," the representative said, the company had removed the page advertising that service from its website. Reddit users did call out a different company for "frequent and sizable online orders": Zoom Books, based in Canada, claims to use "smart logistics" to "acquire, sort, and resell over 2 million books monthly." Booksellers in Germany mentioned similar orders from Zoom, as did a New Zealand bookseller who wrote that Zoom had placed multiple orders, including one for 71 titles. Becker would not confirm or deny whether his bulk buyer was Zoom. On May 26, a book-industry newsletter called Publishers Lunch noted accusations that Zoom was buying books for AI companies. On the following day, the newsletter included a response from Zoom, claiming that it "has no involvement in digitization, scanning, or destruction of books, and we do not sell books for those purposes to our knowledge." (Zoom did not respond to a request for comment.) Some booksellers who received orders from Zoom have reported that the shipping addresses on these orders belong to a company called PrepFort. PrepFort specializes in preparing items for sale on Amazon and Walmart, but as a third-party logistics company, it can route books to other destinations as well. Using such a company would be consistent with the AI industry's practice of acquiring training data through intermediaries, a strategy that has been used for images, paywalled articles, and other media. I emailed PrepFort to see if it could provide any details about these orders. Sufyaan Kalota, a co-founder of the company, responded: "While we would love to assist, we unfortunately cannot share any details regarding our customer agreements due to our confidentiality commitments." In June, the NL Times, an Amsterdam-based publication, reported that some Dutch booksellers had received emails from a Singapore-based organization called 2077AI, which claims to be "revolutionizing AI data," partly by creating training data sets. One bookseller reported that the email from 2077AI was accompanied by a list of 3,000 English-language titles the company wanted to purchase. The examples given were academic books such as Distinct Element Modelling in Geomechanics and Laser Shock Peening of Advanced Ceramics. Such books might legitimately be considered "rare": they had small print runs, contain esoteric knowledge, but are important within a specific field. (404 Media linked to this story but did not reference these books specifically, or 2077AI.) 2077AI did not immediately respond to a request for comment. For now, this is what we know: A lot of used books are suddenly being bought up by companies including Zoom Books. But we do not know how many of these orders are coming from AI companies, or whether these books will be destroyed. Based on the practices of Anthropic and the attempted purchases by 2077AI, it seems very possible that the books are destined for AI companies, but we can't be certain. Regardless, the idea of an AI behemoth destroying books has sent a bolt through the collective psyche. Book destruction is historically linked to authoritarian regimes and attempts to control public discourse, as well as ancient tragedies -- the burning of the Library of Alexandria comes to mind. The encroachment of AI, particularly in the realm of books, is now everyday news: Last month it emerged that a top-selling Kindle book may have been partly written by AI, and grandparents are buying children's books filled with AI-generated slop. Even if a literal book burning is not under way, the process of ingesting millions of books, stripping them of authorship, and blending them into a homogenous "intelligence" branded with the name of a chatbot certainly feels destructive. It masks the hard work and collaboration that goes into knowledge creation, undermines incentives for authors to write books, prevents experts from finding one another, and gives AI companies tremendous power over what information people can access. Given this scale of cultural damage, the physical destruction of books seems almost quaint.
[2]
Booksellers fear their rarest books are being pulped to feed AI. The trail leads to a 'recycler'.
AI firms want old books because they are free of AI slop. Now Australian booksellers fear the rare and irreplaceable ones are being sliced up and pulped to feed the machines. We have written about why AI companies are buying up old books: printed before the internet filled with machine-made text, they are a clean source of language, free of AI slop. This is what that hunger looks like from the other end. In Australia, secondhand booksellers fear someone is slicing apart and pulping the rarest titles on their shelves to train the machines. The fear traces to a US copyright case against Anthropic, brought by book authors. Unsealed files revealed a practice called "destructive scanning": buy the book, slice off the spine, scan the pages, pulp the rest. The Washington Post reported the company tried to keep it quiet. A judge ruled it lawful, as "transformative use." The practice is now understood to be widespread, with 404 Media reporting that intermediaries have stepped in to feed it. That is where Australia's booksellers come in. Several noticed strange waves of orders through the big secondhand marketplaces, Abebooks and Biblio, and started asking who was buying. The buyer, in many cases, was a Canadian "book recycler" called Zoom Books. As the Guardian reported, the orders were price-insensitive and oddly random. Delfina Manor, in Victoria, got three boxes in May and asked them to stop. One order paired a 1970s soil-mechanics manual with a $9 poetry book that had sat unsold for 20 years. 'More than just objects' For sellers drowning in stock, the sales are not all bad. Nick Dawes, who keeps about 500,000 books with his wife Jenny, would happily sell in bulk. His line is the irreplaceable one. "If somebody said to me, look, this is the only known copy, I'm going to cut it up, then I'd say, I think I'd rather not send it to you." Tim White of Books for Cooks draws the same line. A $20 book with half a million copies in Australia, he would not mourn. A one-off, "it's horrific." AI is not the first buyer to destroy books for profit; dealers have long sliced rare volumes to sell the illustration plates. As bookseller Gwenyth Todd puts it, "to me books are more than just objects." Two denials, one silence Both companies named in the story deny the worst of it. An Anthropic spokesperson said it has never bought from Zoom Books, that it sources books from mainstream commercial markets, and that "none of our data acquisition programs buy and destroy rare or antiquarian books." Buying books to train models, it added, is common across the industry. Zoom Books, whose co-founder said on LinkedIn it had been "expanding aggressively," says it resells books intact and does not digitise them, recycling only what cannot be rehomed. It would not say whether it sells intact books on to AI firms for scanning. Its commercial agreements, it said, require confidentiality. That silence is the problem. No vendor has proof their books were destroyed, and the concern is the opacity, not a proven act. A bookseller can screen a buyer they can see. They cannot screen a supply chain that ends behind a confidentiality clause. For a common paperback, no great loss. For the only surviving copy of something, there is no undo.
[3]
Millions of books are reportedly being slaughtered to train AI -- here's what Anthropic told Tom's Guide
The dismantling of classic literature for AI upgrades is disheartening Mention the term "AI" to anyone in the know and it's likely to make them squirm in disgust. Even though the most popular models have proven to be helpful digital assistants for numerous tasks, the companies behind those AI tools are seen as doing more harm than good. Massive AI data centers are harming the environmental stability of the areas they've been housed in, plus the large amounts of memory bandwidth needed to run large language models have led to a global RAM shortage. And then there's the continued serving of "AI slop" that appears in the form of art, music and videos that anyone with a careful eye can recognize. It's well known that AI models are trained by being fed vast amounts of information through a data network -- that data is obtained from web articles, books, academic papers, images, audio and video files. Concerning books, recent reports from The Washington Post and 404 Media have exposed a disturbing trend -- Anthropic (the company behind the Claude AI tool) has reportedly dismantled classic physical literature just so they could make the scanning process go faster in a bid to further train their AI models on their material. Here's an explainer on why and how this trend has come to fruition. Why and how this trend is happening Anthropic's name has been brought up in an operation referred to as "Project Panama." The company reportedly used that codename to define a massive effort to acquire a vast amount of printed books and convert them into AI training data. The AI company is said to have spent millions of dollars acquiring millions of physical books in bulk and proceeding to cut off the spines of said books so the individual pages could be more easily fed through high-speed industrial scanners. This process then led to those pages being converted into machine-readable text for machine learning -- that information eventually became a part of a larger dataset used to train Claude AI. And due to those original books' bindings being removed, they were disposed of as they were now considered nothing but trash. The Washington Post's report on this whole matter featured a quote from an internal planning document connected to Anthropic's ultimate goal. "Project Panama is our effort to destructively scan all the books in the world," that document stated. "We don't want it to be known that we are working on this." Clearly, word got out about the AI giant's destruction of so much classic literature -- this reveal resulted in a class action lawsuit brought against them by a collective of authors whose pirated work had been used to train Claude. This past July, Anthropic settled with those same authors at $1.5 billion. Reuters reported that the court overseeing that case ruled that, while training AI on books is classified as fair use under copyright law, Anthropic violated the copyrights of several authors and their publishers by keeping 7 million pirated books of theirs in a central library. While companies such as OpenAI, Google and Meta have been embroiled in legal trouble regarding the use of books being used for AI training, Anthropic is the only one that's been exposed thus far for actually dismantling physical books to train its models. Another report from 404 Media alluded to a company called ISBNdb, a commercial book metadata and sourcing service that helps businesses locate books and bibliographic information, being used to find and purchase large quantities of physical books for sale to AI companies for their training efforts. Since that report went public, ISBNdb has scrubbed its webpages of any info related to training AI. Additionally, the company has denied ever buying, scanning or selling any books for AI training purposes and noted that its site was merely a "test of market interest." Here's what Anthropic had to say A spokesperson for Anthropic tells Tom's Guide, "Claude is trained on a mix of publicly available web data, commercially acquired datasets, and data we generate ourselves," the Anthropic spokesperson told us. "Sourcing books is a widely used approach for training large language models across the AI industry. None of our data acquisition programs buy and destroy rare or antiquarian books." Anthropic also provided more background on how they engage in the practice of using books to train AI: * Sourcing books is a widely used approach for training LLMs across the AI industry. * None of our data acquisition programs buy and destroy 'rare' or 'antiquarian' books. We buy books from regular commercial markets. * Much of the chatter online seems to reference our recent Bartz settlement; as a reminder: We settled this case last summer, after the court found that training AI on books is allowed under copyright law [known as 'fair use'] -- a ruling that still stands. * As Judge Alsup wrote: "Like any reader aspiring to be a writer, Anthropic's LLMs trained upon works not to race ahead and replicate or supplant them -- but to turn a hard corner and create something different." * Bartz specifically concerns claims that Anthropic improperly downloaded materials from two online libraries -- the LibGen and PiLiMi datasets. * Anthropic never commercially released any model trained on the LibGen or PiLiMi datasets, and earned no revenue from any such model. * We are pleased that more than 91% of authors and publishers covered by the settlement have claimed their share of the payment. Bottom line The visual of legendary books having their spines cut off, turned into nothing more than AI training data and later discarded is the stuff nightmares are made of. Rare and even out-of-print books may have been obtained by Anthropic for their past efforts, but we might never truly know until years later if the names of those demolished books become widely known. I'm sure no one wants to live in a world dominated by AI tools that have been trained on reading material that's not readily available in their original physical form. Follow Tom's Guide on Google News and add us as a preferred source to get our up-to-date news, analysis, and reviews in your feeds. Subscribe to Tom's Guide on YouTube and follow us on TikTok.
[4]
Why is Anthropic destroying books? | Kathryn James
The AI company apparently found destructively scanning 'all the books in the world' easier than dealing with copyright in its quest for training data Should we destroy all the books in the world? An answer to this question can be found in the court documents of Bartz v Anthropic PBC. The northern California district court case, decided in late July this year, highlighted the improbably named "Project Panama", one of the AI company Anthropic's efforts to improve its large language model Claude. "What is Project Panama?" court exhibit 21 asks, in an internal memo. The answer: "Project Panama is our effort to destructively scan all the books in the world." The memo advises discretion: "Why use a codename? ... [B]ecause we don't want it to be known that we are working on this." Destructive scanning was Anthropic's solution to a problem: in order to "train" Claude, Anthropic had to procure a large, high-quality language dataset, preferably one created before 2022 and the corrupting influence of generative AI on contemporary text. Claude needed as many language combinations as possible, to improve its ability to predict language outcomes. Anthropic needed data, lots of it, of very high quality. Books, as it happens, remain one of the best sources for complex, high-quality, long-form text. As the court's decision relates, Anthropic hoped that books' "well-curated facts, well-organized analyses, and captivating fictional narratives" would help "Claude write as accurately and as compellingly as Authors". Anthropic had a choice: it could have secured copyright permission to use existing e-books. This would have required the "legal/practice/business slog", as Anthropic's co-founder and CEO phrased it, of managing copyright. Rather than engage with the texts' owners, Anthropic first chose to use pirated sources instead, a decision informing the company's $1.5bn out-of-court settlement with authors. When that approach seemed too complicated (or, as the court decision phrases: "Anthropic became 'not so gung ho about' training on pirated books 'for legal reasons'"), Anthropic turned to destructive scanning. As it happened, the judge ruled that using proprietary material to "train" an LLM did not, in and of itself, constitute an infringement of copyright. To the court, it would seem, "training" a corporate product is equivalent to training any human, teaching how to read in order to learn how to write. Should we be surprised that destroying printed texts seemed easier to Anthropic than working with their human authors? Either way, the decision to use destructive scanning turned the issue into one of logistics. Under US copyright law, the "fair use" doctrine allows you to make "transformative" use of copyrighted works without the owner's permission. Anthropic took printed books and scanned them, "transforming" or remediating them into a new, electronic format. They then disposed of the original printed copy: the "destructive" part of destructive scanning. Along the way, Anthropic's vendors had already sliced the spines and edges of the books, to scan them more easily before destroying them. "One replaced the other," as Judge William Alsup wrote, noting: "There is no evidence that the new, digital copy was shown, shared, or sold outside the company." Anthropic hired an experienced logistics manager, sourced the books from vendors, hired staff and housed the books in a warehouse, and found a digitization vendor to take on the project of the destructive scanning itself. The case exhibits show warehouses of books neatly stacked and labelled on shelves, staff moving between them. As the court's decision relates, Anthropic's vendors "stripped the books from their bindings, cut their pages to size, and scanned the books into digital form - discarding the paper originals". In the court's images, stacks of books await destructive scanning. Not seen: the clean-up project of shredding and disposing of the books (or, "all the books in the world", reformatted as recycling or landfill). Bookishness, or the range of meanings attached to the book as cultural object, has always lived alongside the book's role as textual instrument. Witness the example of images of Donald Trump, holding up a copy of the Bible at St John's during the protests of 1 June 2020. Yet unlike many other countries, the US has very few legal provisions relating to the regulation or export of American cultural heritage, and few to none governing the treatment of books. In the context of the court's decision in Bartz v Anthropic PBC, destructive scanning is an effective mechanism to strip authorial involvement from printed texts, in order to use the content of those works to improve the functions of an LLM. It is both legal and less regulated than strip mining. What are the consequences if, as seems likely, this practice is adopted by other generative AI companies, now and in the future? How many warehouses of destructively scanned books would be too many? There is no endangered list for printed works, and little regulation of what might constitute survival of the rare or unique. Still further, there is no formal understanding of the human-generated textual object, in and of itself, as a category of cultural asset or heritage that might require protection. If we take seriously the 2022 threshold, as a moment when AI-generated text began to make significant entry into the textual record, should we start to think of the "wholly human author" as an emergent category of collections preservation and stewardship? There is more to say here (what to make, for instance, of the "forever" research library Anthropic states it intends to create), but let me close by observing that Bartz v Anthropic PBC is also a powerful statement on the importance of books or long-form text - and of readers. Judge William Alsup noted: "For centuries, we have read and re-read books. We have admired, memorized, and internalized their sweeping themes, their substantive points, and their stylistic solutions to recurring writing problems." We should worry that Anthropic decided it was easier to scan and destroy physical books than to deal with the "legal/practice/business slog". We should worry that the current understanding of fair use allowed Anthropic to decide that it was easier to buy and destroy "all the books in the world" than to pay the creators of those works. But perhaps the most telling aspect of this case is that it was so important to Anthropic to have access to a dataset of complex, long-form, uncorrupted text. Let's ask ourselves why that text should seem so profitable. One response to Bartz v Anthropic PBC might be to refuse to devalue our lives as readers and writers, to claim ownership of the cultural spaces in which our thought and words are created, shared and preserved. The risk with generative AI, as this single instance with Anthropic indicates, is that we cede the means of production of our large language lives: that we turn from creators to consumers, and hand the generative promise of our work to large language models and their proprietors.
[5]
Millions of books face destruction as AI companies race for training data
AI is now killing centuries-old book volumes that survived wars * ISBNdb ships up to a million books anonymously to AI labs * Pre-2022 books are prized because chatbots never touched their text * Spine-cutting scanners destroy originals to speed up digitization for training AI companies are increasingly turning to printed books published before 2022 as preferred training material because those works predate the widespread use of AI-generated content. Large-scale scanning operations reportedly involve cutting book spines, separating pages, and destroying physical copies to create digital datasets for large language models. The practice has attracted growing criticism because some books entering these pipelines are reportedly extremely rare, raising concerns about irreversible cultural losses. Pre-2022 books become valuable AI training material Reports by 404 Media found data broker ISBNdb supplies physical books in bulk to AI developers seeking human-written material unaffected by modern chatbot output. The company argues books published before 2022 offer cleaner datasets because they cannot contain text generated by contemporary large language models. They are often considered "dense, edited, authoritative," in contrast to internet content increasingly filled with machine-generated material of uncertain quality. The approach also attempts to avoid so-called model collapse, in which AI systems gradually lose quality after repeatedly training on synthetic content generated by earlier models. ISBNdb additionally argues that older printed works avoid deliberate data-poisoning techniques authors increasingly use to disrupt AI training through carefully modified documents. However, there are reports that many of these books are scanned using high-speed equipment. This equipment requires workers to remove the spine before feeding individual pages through automated imaging machines. That process reportedly destroys the original volume, making rapid digitisation considerably cheaper than slower preservation methods designed to keep books physically intact. Secrecy and legal rulings fuel preservation concerns ISBNdb openly acknowledges reputational concerns surrounding the practice while offering strict non-disclosure agreements that keep customer identities confidential throughout commercial engagements. Its website reportedly states, "'AI company destroys two million books' is not a headline that generates sympathy," while suggesting clients describe the process as digital preservation. Such a level of destruction is an order of magnitude bigger than the loss of the Library of Alexandria. Yet, it is unfolding with none of the outrage that history reserves for burned libraries. Booksellers interviewed by 404 Media said some volumes entering these scanning programmes have very few surviving copies after enduring wars, fires, and centuries of handling. Critics argue that unlike websites or widely available modern publications, exceptionally scarce historical works cannot simply be reproduced after their physical copies disappear forever. A recent United States court ruling involving Anthropic found that scanning legally purchased books for AI training constituted fair use under specific circumstances. Part of that reasoning held that destroying each printed copy during scanning meant one legal copy effectively replaced another rather than creating multiple copies. In response to a critic (@Hedgie) of this method on X, Elon Musk said, "I've asked the SpaceXAI team to preserve any rare books in a library and scan them the hard way," suggesting an alternative approach. If significant awareness is not created, this quiet erasure of irreplaceable books risks becoming the defining act of cultural loss for this era, remembered only after it can no longer be undone. Follow TechRadar on Google News and add us as a preferred source to get our expert news, reviews, and opinion in your feeds.
[6]
AI companies are buying and destroying old books for training data
AI companies have spent years pulling training material from the internet. Now, as more of the web fills with AI-generated writing, some are looking for text in a place largely untouched by chatbots: used-book shelves. Until July 28, ISBNdb, a company that operates an online book database, advertised its ability to source as many as 1 million physical books per order for AI developers. The since-deleted page offered books "tailored to your LLM training needs," according to 404 Media, including older, specialized, rare, and out-of-print titles gathered from used bookstores and other catalogs. The company also promised confidentiality. Its marketing materials advertised a "strict NDA on every engagement" and said clients' identities, strategies, and acquisition targets would not be disclosed. That secrecy attracted attention because converting physical books into AI training data can involve destructive scanning: cutting off their bindings, feeding the loose pages through industrial scanners, and discarding or recycling the originals. By July 28, ISBNdb had removed the sourcing page and its promise of an NDA. The company said the service had been part of "exploring demand" and that it had "chosen to pivot away from that direction," in a news update. Still, ISBNdb's brief sales pitch offered a glimpse into a book-buying operation that one major AI developer has already carried out on a much larger scale. Federal court records show that Anthropic purchased and scanned millions of physical books, while used-book sellers in the United States and Europe have recently reported unusual bulk orders for obscure and out-of-print titles. The reports have opened several practical questions: Why have older books become so valuable to AI developers, how widespread is destructive scanning, and what happens to the originals once their pages become training data? Why AI companies are hunting for old books Older books offer something that has become surprisingly difficult to guarantee online: writing produced entirely by humans. Large language models are trained on enormous quantities of text collected from websites, articles, books, code repositories, and other digital sources. Since generative AI tools became widely available, however, the internet has filled with machine-written summaries, product listings, social posts, and articles. That creates a problem for developers assembling new training datasets. If a model is trained too heavily on material produced by earlier models, it can begin reinforcing their mistakes while losing some of the variety and less common information contained in the original human data. Researchers call the process "model collapse." A physical book printed before the current AI boom offers a relatively clean alternative. Unlike a webpage that may have been quietly generated or rewritten by a chatbot, an older book provides a more dependable record of human writing. Books are also edited, structured, and often contain specialized information that cannot easily be found elsewhere online. ISBNdb leaned heavily on those qualities in its sales pitch. "Books represent curated, peer-reviewed, domain-specific human knowledge, structured in a way no web crawl can replicate," the company wrote. "Dense, edited, authoritative." ISBNdb also emphasized that many potentially useful titles had never been fully digitized. Its marketing effectively presented the pre-chatbot bookshelf as a reserve of human-created training material that had yet to be mixed with AI output. That demand did not appear out of nowhere. It follows a much longer effort to turn printed books into searchable digital information. AI didn't invent mass book scanning People have been converting books into digital files for decades. Project Gutenberg began creating electronic versions of public-domain works in 1971 and now offers more than 75,000 free ebooks. The industrial-scale version arrived with Google Books in 2004. Working with libraries and publishers, Google scanned books and made their contents searchable online. By 2019, the company said it had assembled a collection of more than 40 million books in over 400 languages. Those scans also helped form HathiTrust, a digital repository created by research libraries in 2008 to preserve their collections and make as much of the material publicly accessible as copyright law allowed. The project prompted an earlier version of today's copyright fight. Authors sued Google for scanning copyrighted books without permission, but a federal appeals court ruled in 2015 that the project qualified as fair use. The court found that creating a searchable index and displaying limited snippets gave the books a new purpose without providing readers with a replacement for the originals. Other digitization programs have emphasized preservation and access. The Internet Archive, for example, says it has digitized more than 25 million books since 2006 using nondestructive scanning designed to keep the bound volumes intact. But converting a book into data does not necessarily make it publicly accessible. That distinction concerned Charlie D. Becker, a second-generation bookseller whose family runs Becker's Books in Houston. His store recently received a single order for 70 obscure titles, including a 1995 guide to metropolitan Denver and manuals explaining how to use WordPerfect in 1991. After investigating, Becker suspected the buyer was using an algorithm to find underpriced books and relist them on Amazon, rather than acquiring them for AI training. Most of the titles were either unavailable on Amazon or listed there for as much as 20 times his store's price, and the shipments appeared to be going to Fulfillment by Amazon preparation companies. Still, Becker said an obscure book can be especially easy to lose because few people consider it worth preserving. Books that fail to sell through Amazon's fulfillment system may eventually be liquidated, while some titles have little more than scattered listings across private databases to prove they existed. "Everyone assumes the internet preserved everything," Becker wrote in a longer account of the orders. "It didn't." Anthropic has already scanned millions of books The buyers behind the recent bookstore orders remain unclear. The destructive scanning process, however, is no longer hypothetical. In 2024, Anthropic hired Tom Turvey, a former Google executive who had worked on partnerships for the Google Books project. According to a June 2025 federal court ruling, Turvey was tasked with helping the company obtain "all the books in the world" for an internal research library. Turvey initially contacted publishers about licensing their books, but those conversations did not continue. His team then approached major book distributors and retailers about purchasing print copies in bulk. Anthropic ultimately spent millions of dollars acquiring millions of physical books, many of them used. Service providers removed the books from their bindings, cut their pages to the appropriate size, and fed them through scanners. The searchable PDF files went into Anthropic's internal library. The paper originals were discarded. Engineers could then select groups of those books for inclusion in datasets used to train the large language models behind Claude. That process became central to a legal battle over how Anthropic obtained its training material. In June 2025, U.S. District Judge William Alsup ruled that the company's use of books to train Claude was transformative and qualified as fair use under the specific circumstances of the case. Alsup also found that converting legally purchased print books into digital files could qualify as fair use because Anthropic destroyed each physical copy and replaced it with one internal digital copy. In other words, the company did not keep both versions. The ruling did not give AI companies blanket permission to copy any book they could find. Alsup drew a sharp distinction between the print books Anthropic had legally purchased and the more than 7 million pirated books the company had downloaded and stored in a permanent digital library. Purchasing physical copies of some titles later did not erase the original piracy, he found. That distinction eventually became expensive. On July 20, a federal judge approved Anthropic's $1.5 billion settlement with authors and publishers, resolving claims involving approximately 482,000 pirated books. Eligible rights holders are expected to receive about $3,000 per title, according to Reuters. For some authors, the payment does not resolve the larger disagreement. Charles Graeber, one of the case's original plaintiffs, told NPR that he was proud authors had secured a substantial settlement, but said the case had cost him more than two years of time, travel, and professional opportunities. Fellow plaintiff Andrea Bartz questioned a system that allows companies to train commercial models on legally purchased books without negotiating separate licenses with the people who wrote them. "The algorithm is being used to essentially try to put us out of a job," she told NPR. Anthropic has maintained that training AI models on books is protected by fair use. The court's ruling nevertheless helps explain why physical books may be especially attractive to developers: A lawfully purchased copy gives the company a much stronger legal position than a file downloaded from a pirate library. Destroying the original may be a legally useful distinction. For many readers, it is also the most unsettling part of the story. Who is buying all these books? Anthropic's operation is documented in court records. The source of the more recent bookstore orders is harder to pin down. One bookseller specializing in uncommon and low-circulation titles told 404 Media that his weekly sales jumped from roughly 20 books during a good week to several hundred after the orders began arriving in April. The requested titles did not appear to share a subject, author, genre, or language. They did, however, all have International Standard Book Numbers, or ISBNs, leading the seller to suspect that they had been selected through a book database. The surge was financially helpful and allowed him to clear inventory that might otherwise have remained unsold. He was less enthusiastic about where the books might be going. "I don't like the end-use, and I don't like that uncommon books are being pulped," he said. Booksellers in Europe have reported similarly broad requests. An antiquarian bookseller in the Netherlands received a list of approximately 3,000 English-language books organized by ISBN, ranging from an academic study of Irish folklore to a technical book about laser shock peening. There is no public confirmation that every unusual bulk order came from an AI company or that every book purchased through these orders was destroyed. There is also no evidence that developers are intentionally hunting for the final surviving copies of rare titles. The uncertainty itself is part of the concern. A company purchasing from a massive ISBN list could sweep up uncommon or out-of-print editions without first checking how many physical copies remain. Why the story struck a nerve Once the reports reached social media, they were accompanied by an unsettling visual: a cutting blade moving inch by inch through a book's spine. Comparisons to Fahrenheit 451 and the Library of Alexandria followed quickly. Book destruction carries a particular weight because it has historically represented more than the loss of paper. The 1933 Nazi book burnings targeted works deemed "un-German," including books by Jewish, pacifist, and left-wing writers. The act became an enduring symbol of censorship and the suppression of ideas, according to the United States Holocaust Memorial Museum. That symbolism is central to Fahrenheit 451, Ray Bradbury's 1953 novel about a society where books are outlawed and burned. Bradbury said his warning extended beyond government censorship to television reducing knowledge to digestible fragments and eroding interest in reading. Users have also invoked the Library of Alexandria, another enduring symbol of lost knowledge, although historians believe it declined gradually through political upheaval, reduced support, neglect, and repeated damage rather than disappearing in one catastrophic fire. Elon Musk joined the conversation on July 27. He wrote on X that he had asked the SpaceXAI team to preserve rare books in a library and scan them "the hard way," without removing their spines. As the discussion spread, users resurfaced a 2011 Cracked article about libraries, universities, and retailers destroying unwanted books years before the current AI boom. That history has informed a less alarmed response. Some users argued that the books shown in warehouses appeared to be ordinary, unwanted inventory rather than irreplaceable artifacts. If a book would otherwise be recycled without being read again, they asked, could scanning it first preserve something that would have been lost? AI adds a complication. The text may survive, but inside a private company's research library rather than a public archive. A forgotten manual or travel guide may have little resale value, yet still contain exactly what an AI developer wants: edited human writing created before the flood of chatbot output. That has led authors and publishers to argue that they should have more control over how their work is used, especially when it is helping companies build commercial products. AI developers, meanwhile, continue to argue that training a model is a transformative use of the material rather than a replacement for the original books. ISBNdb has removed its sourcing page, but the demand behind it remains. The web's AI problem has sent developers back to the bookshelf. The next chapter will depend on whether they can extract what they need without leaving those shelves any emptier.
[7]
This Dutch bookseller thought a request for 3,000 copies was 'spam or phishing.' Instead, AI companies are scanning and destroying books to train AI | Fortune
It was just another sunny summer day in Haarlem, Netherlands, when an unusual request landed in Dutch antiquarian bookseller Pieter de Vries' inbox. A woman named Natalia, identifying herself as being with a company called "2077AI," was looking to source what she described as a "fairly large order" of books. Attached was a spreadsheet listing more than 3,000 ISBNs along with instructions to match titles, prepare quotes and estimate shipping costs directed to China. De Vries barely looked at it, dismissing the message as spam or phishing without reading past the first few lines. It wasn't until weeks later, when a Dutch journalist investigating the procurement request contacted him, that he learned of the request's connection to artificial intelligence. "I was shocked!" said de Vries, who provided Fortune with the email and subsequent spreadsheet containing 3,001 titles, mostly published between 2020 and 2021 by academic publishers like Emerald Publishing, Elsevier, Wiley, Routledge and Oxford University Press, and ranging from business and education to engineering, public policy and medicine. "Very unusual," de Vries said. "That's why I regarded the request as spam or phishing." De Vries wasn't the only one. He told Fortune several other Dutch antiquarian booksellers received the same request and, like him, dismissed it. (They didn't take "it seriously" and "considered it as scam," he said). While the inquiry had little effect on his business, he said, it could have been a different story for booksellers who specialize in selling secondhand books at scale. "I deal in old and rare books," de Vries said, adding he rarely receives book orders by e-mail and if he does it's at most for only one book. "These books are expensive and exclusive. So I do not do bulk trade." AI's appetite for books The email De Vries received never mentioned artificial intelligence or explained what the books would be used for., but it surfaced as AI companies increasingly look beyond the open internet for high-quality written material to train large language models. Last summer, public attention to AI companies' use of books intensified when court records revealed Anthropic had purchased millions of physical books, removed their bindings, scanned them and discarded the originals to build a searchable digital library used to train its AI models in what internal documents called "Project Panama," according to The Washington Post. The process, known as "destructive scanning," involves cutting the spine from a book so its pages can be fed through high-speed scanners before the remaining physical copy is discarded. The case later settled after a federal judge ruled that using legally purchased books to train AI models constituted fair use under copyright law. Separate claims involving Anthropic's downloading of books from the LibGen and PiLiMi online libraries were resolved through the settlement. According to an Anthropic spokesperson, Claude models are trained on a mix of publicly available web data, commercially acquired datasets and internally generated data. Anthropic purchases books through regular commercial markets, and none of its data acquisition programs acquire or destroy rare or antiquarian books. The hidden supply chain The court records showed what happened after books were acquired, not how companies source millions of physical volumes in the first place. Earlier this year, 404 Media reported that ISBNdb, a company best known for maintaining metadata tied to International Standard Book Numbers (ISBNs), the unique numerical identifiers encoded in the barcodes found on most books, had begun advertising services to help AI labs source bulk printed book purchases ranging from 1,000 to 1 million books. According to the report, the company's website said the service was tailored to "LLM training needs" and "delivered at the scale AI demands." The webpages have since been removed. ISBNdb later told Fortune the service was never launched, it never purchased or scanned books for AI companies and the webpages reflected an exploratory concept rather than an active business. Similar reports have emerged in Germany and Switzerland, where local media reported that secondhand booksellers received unusual bulk orders for highly specialized books that appeared inconsistent with traditional collecting or resale. Zoom Books, a Canadian company identified in some of the reports, denied these allegations and told Swiss broadcaster SRF the purchases were part of its regular recycling and trading model. For an industry built on algorithms, the email offers a rare glimpse into something much more tangible: the physical supply chain of books that may ultimately power AI models. Fortune contacted 2077AI seeking comment about the purpose of the procurement request and whether the books were intended for AI training. The company did not respond to requests for comment before publication.
[8]
Project Panama: How Anthropic shredded books to train its AI chatbot
Recently unsealed court documents reveal how Anthropic secretly bought, destroyed and scanned books to train its AI chatbot, Claude. It sounds like a Ray Bradbury story come to life. Somewhere in America's sprawl of warehouses, book spines were sliced by machines like cuts of salami. Scanners then processed their contents in bulk. Within milliseconds, work that may have taken humans years to produce was uploaded into the digital ether, and later used to train AI models to replicate the voice and structure of our languages. From there, the paper became waste, which was shipped off to recycling plants to be pulped and repurposed into products like toilet paper and cardboard boxes. In a name that evokes both a covert operation and documents exposing financial sleight of hand, Project Panama was Anthropic's secret plan "to destructively scan all the books in the world," an internal planning document unsealed in legal filings revealed in late July. "We don't want it to be known that we are working on this," the document said. The revelation of Anthropic's shadowy efforts - first reported by the Washington Post - has cast another pall on the AI arms race. It has also offered one of the clearest looks yet at how one of the world's biggest AI firms views the work of authors, artists, photographers and other creators. How Project Panama fed the paper mill Project Panama began in early 2024, as Anthropic executives sought books predating those available online that could teach its AI chatbot Claude "how to write well" instead of mimicking "low quality internet speak", according to newly released documents. The startup purchased tens of thousands of books at a time from multiple vendors, including Better World Books and UK-based World of Books. It then outsourced the heavy work of scanning and destroying books to a vendor, which used a hydraulic machine to cut off the spines and slice pages one by one so they could be quickly copied and uploaded using industrial imaging equipment. According to a vendor that ultimately worked with Anthropic, the company was "seeking an experienced document scanning services vendor to convert from 500,000 to two million books over a six-month period." Why books instead of the internet? Internal planning documents suggest Anthropic executives believed the initiative would lead to higher quality communication and address AI model collapse - a kind of degradation that occurs when AI trains on text contaminated by its own output. Unlike much of the modern web, books offer vast quantities of human-written and professionally edited literature yet to be spoiled by AI-generated content. The challenge for the AI companies was a matter of how to obtain them. But the court filings suggest that Anthropic executives may have believed they could get away with it, and even if they couldn't, that the end would justify the means. According to one filing, Anthropic co-founder Ben Mann acquired a trove of copyright-protected content from a "shadow library" called LibGen in June 2021. The filing even included a screenshot of Mann's web browser in the process of downloading pirated books. The following year, Mann shared a link to the Pirate Library Mirror - a full catalogue of illegally acquired books - with other Anthropic employees. "just in time!!!" he wrote. The legal reckoning Project Panama only became public through thousands of pages of court documents unsealed in a copyright lawsuit brought by a group of authors, who alleged Anthropic illegally used their work to train Claude. In 2024, a group of authors, led by novelist Andrea Bartz and nonfiction writers Charles Graeber and Kirk Wallace Johnson, hit Anthropic with a class action lawsuit, alleging the company violated copyright laws by using pirated books to train its AI model. Last month, Anthropic - currently valued at US$965 billion (€875 billion) - agreed to pay $1.5 billion (about €1.3 billion) to settle the case, covering about 500,000 eligible works. "It is the largest known copyright recovery in history," said Justin Nelson, the authors' lead attorney. The case is the first of a wave of them against AI companies to settle. Meta, OpenAI, Google and Microsoft are also facing copyright lawsuits from authors making similar allegations of piracy. Several authors who opted out of the class action settlement have individual cases pending against Anthropic, too. What the ruling means As other cases work their way through the courts, the settlement last week may be only the first chapter in a much longer legal fight. In the original ruling last year, a judge found that the company's use of lawfully purchased books was protected as fair use under US law. However, the ruling drew a distinction between books the company bought and scanned and the millions of books it had previously pirated. "The court's landmark June 2025 ruling remains intact," Anthropic deputy general counsel Aparna Sridhar said in a statement to the Post. "The issue we settled on was about how some materials were acquired, not whether we could use them to develop [AI]." For some, however, the unsealed documents may have delivered the more lasting verdict, shattering whatever illusion may have remained that the world's biggest tech firms understand cultural value. Let alone value the sanctity of human creation.
[9]
Company that said it could scan and destroy books for AI data-harvesting has deleted that part of its website: 'no such service was ever brought to life'
Remember the story about AI companies destroying millions of copies of books while scanning them to use as training data? In terms of bad optics, "destroying books" is right up there with kicking puppies, which is why the AI companies relied on third-party rebuyers to obfuscate what they were actually doing. It didn't work, and now as 404 Media reports (having broken the original story), third-party book database company ISBNdb has deleted the part of their website advertising that they could be "your streamlined partner for sourcing printed books in bulk, tailored to your LLM training needs, delivered at the scale AI demands." That page is still available on the Internet Archive, of course, as is their blog post about "Reframing the Destruction Narrative" in which they make wild claims like: "The book is not destroyed. Its value has migrated. The paper returns to the material cycle; the knowledge enters the intellectual one." The idea that asking a chatbot to hallucinate a summary of a book has exactly the same value as actually reading it seems bugloving nuts to me, but what do I know? I'm not getting my knowledge from the intellectual cycle. Anyway, ISBNdb is now claiming it never actually offered the service it dedicated multiple parts of its website to, telling 404 Media, "The page was a test of market interest; no such service was ever brought to life." Somebody out there sure is buying up a lot of old and obscure books, however. As reported by Guardian Australia, multiple secondhand booksellers have seen an uptick in online sales from third-party rebuyers, describing them as "waves of orders that did not fit with usual customer patterns."
[10]
'More than just objects': Australian book sellers raise alarm over 'horrific' destruction of rare titles to feed AI
Secondhand booksellers believe they may have been caught up in the AI supply chain that sees old books scanned then destroyed "Sometimes books tell you more than the content," says Tim White, as he lines up carefully dust-jacketed books and pamphlets in plastic sleeves on temporary shelves in Melbourne University's Wilson Hall. He's preparing his stall for this weekend's Melbourne Rare Book Fair. For people like White, who runs independent bookshop Books for Cooks, the book as a physical object is just as important as what's inside it. "We're interested in where food and society connect ... So we just keep pushing the boundaries, looking for things that show unusual aspects of that, things that tell us a bit more about what it is to be human," he says. "That's what most of us [rare booksellers] do. We handle things that have a story." But vendors and rare books experts are increasingly worried that these artefacts may be under threat of destruction by generative AI companies seeking to feed their ravenous technology with new texts. Fears stem from a copyright lawsuit brought against AI company Anthropic in the United States by book authors. Unsealed court files revealed earlier this year that the company was engaging in the practice of "destructive scanning" - buying physical books, slicing off the spines to more efficiently scan the pages, then pulping the remains - and had sought to keep it quiet. The judge ruled this was lawful under US copyright law, as it amounted to "transformative use". The practice is now understood to be widespread among generative AI companies, and recent reports suggest intermediaries are stepping up to facilitate the process. Australian secondhand booksellers are wondering if they have been caught up in the AI supply chain after they recently experienced waves of orders that did not fit with usual customer patterns. Multiple vendors confirmed to Guardian Australia they had received orders from a Canadian company called Zoom Books, which describes itself as a book recycler. The orders usually came through large secondhand marketplaces like Abebooks or Biblio, where inventory can range from cheap paperbacks to rare, expensive items. Delfina Manor, who runs Good Reading Secondhand Books in Benalla, Victoria, said she received orders from Zoom Books in May amounting to three full boxes, but had to write to the company and ask them not to order in such volume again. "They ordered, paid in advance, and they didn't quibble over the postage ... [but] it would have been about 30-40 books, and so finding them, packing them, making sure you hadn't missed one out ... It was just driving me nuts," Manor said. John Sainsbury, owner of Melbourne's Sainsburys Books, confirmed a slew of orders from Zoom Books had come through the store's online listings around the same time. The orders were price insensitive, and appeared entirely random - niche titles, "bottom-end" and sometimes decades-old stock. The orders resulted in Sainsburys shipping multiple boxes overseas. Sainsbury described his company as "a general bookseller at an antiquarian book fair", as he doesn't just deal in rare material. The profile fits the pattern of vendors who have received similar orders - high-end specialist sellers did not appear to have been approached. Some vendors speculated the company may be buying the books on arbitrage - common enough in the secondhand book trade - but the specific requests nevertheless seemed odd. In one order, a 1970s manual on soil mechanics was bought alongside a 2010 book called Born to Thunder: Champions of New Zealand Cycling, a collection of Early Australian Poetry (1982), and a local history of the Melbourne suburb of Hawthorn. The poetry book sold for $9 and had been on the shelf for 20 years. White says the issue raises "an interesting tension". "If it's a $20 book of which there's ... 500,000 copies in Australia, I wouldn't cry," he says. "I'd have different issues around copyright." "If it's a book that has that storytelling element to it, or it's a one-off, it's horrific." For booksellers struggling to catalogue their vast stock of books, let alone sell them, such sales aren't necessarily bad. Nick Dawes, who runs Grant's Bookshop with his wife, Jenny, estimates they have about 500,000 books in their warehouses. "If somebody wanted to come and buy all my books and cut them up, well ... I'd be unhappy at one level, but I'd be happy at another," says Dawes. "But if somebody said to me, look, this is the only known copy, I'm going to cut it up, then I'd say, I think I'd rather not send it to you." AI companies are not the first to destroy books for profit. Some vendors source rare books in order to slice them up and sell the illustration plates individually. "When I sell a book and then find the prints available individually, I just find it physically abhorrent," says rare bookseller Gwenyth Todd, from Chatelaine Books. She screens buyers accordingly, so the books don't go "to the wrong home". "To me books are more than just objects." An Anthropic spokesperson said the company has never bought from Zoom Books, but that "sourcing books is a widely used approach for training large language models across the AI industry". Anthropic procures its books from mainstream commercial markets, the spokesperson said. "None of our data acquisition programs buy and destroy rare or antiquarian books." Zoom Books' co-founder, Manroop Gill, said in a recent LinkedIn post that the company had been "expanding aggressively" in the past months. Zoom Books has previously denied it destructively scans books. A spokesperson told Guardian Australia: "We acquire secondhand books and resell them intact. We do not digitize books. Our priority is always reuse, and when a book can no longer be rehomed, it is responsibly recycled." The company did not disclose if it sold intact books to AI companies for destructive scanning, saying: "Our commercial agreements require confidentiality, so we do not disclose or discuss our customers."
[11]
AI companies are reportedly buying up and destroying millions of old books for training
According to a new report from 404 Media, AI companies are not only using books to build and train powerful AI models, but they're also destroying them. Yes, to digitize millions of books in a way that saves them time and money, AI companies are removing spines and individual pages to make the whole process faster. It's a system referred to as 'destructive scanning,' where the source material is effectively destroyed. This process came to light a year ago, when Anthropic's "Project Panama" initiative was taken to court in the United States for this very reason. Even though the court ruled in favor of Anthropic because it legally purchased these books to turn them into digital copies, it led to widespread backlash because it meant that millions of books were destroyed, everything from paperbacks to rare prints and editions. Even though destructive scanning is viewed as 'fair use' by federal courts, AI companies are now looking to shield themselves from criticism by working with intermediaries to acquire books at scale on their behalf. The 404 Media report notes that these intermediaries promise confidentiality as they source physical copies of books in vast quantities, in the range of hundreds of thousands, or more, at a time. One of these companies is ISBNdb, which boasted that it could source rare books from "library shelves, used bookstores, and out-of-print catalogs," with the option to buy 1 million books at a time with a "strict NDA" to protect AI companies from being exposed. That's 'boasted' in the past tense because this has since been taken down, in a move to minimize the fact that it's now making a lot of money from its "print books for AI training" business. But not before it was archived. Based on the revelations, it certainly sounds like AI companies are trying to buy up every single book, digitizing them, and then destroying them. This isn't alarmist, but reality. And the reason it's happening, and reportedly accelerating, is that we're living in a post-generative AI world. The value of human-created works, especially in the literary space, is increasing. Untainted by "AI slop," these books that are free from any AI influence are viewed as extremely high-quality training data because they're entirely written by humans.
[12]
Suppliers Scramble As Rare Books Destroyed To Train AI
ISBNdb says their controversial books-to-AI pipeline landing page was 'a test of market interest' after disgust of a program straight out of Fahrenheit 451 This week, heavy readers and casual posters alike were aghast at emerging reports of rare books being mulched in bulk as a sacrifice for LLMs. Word traveled that book suppliers were receiving abnormally large orders, believing that AI firms sought cleaner source material while skirting copyright laws. At the center of the controversy was ISBNdb, a book database who not only encouraged the trend but seemed to facilitate it. The optics bit back, as the site now tries to distance itself from the controversy, scrubbing its own posts on the subject. "We've seen the recent coverage about a marketing landing page on our site, and we understand the concern it raised," writes ISBNdb in an update. "We don't train AI models, and we never have. The page was a test of market interest; no such service was ever brought to life. We've taken the page down." A few days earlier, 404 Media reported about the concerning trend of suppliers suddenly hit with bulk book orders. With schools and libraries lean on resources, the decreased business has left them vulnerable. It was suspected that AI firms were making these new orders, but putting themselves in front of the crosshairs was ISBNdb, a database who seemingly introduced a service to patch AI firms through to depositories. "The world's best AI training data is sitting on a shelf," read the now deleted landing page. "Books represent curated, peer-reviewed, domain-specific human knowledge, structured in a way no web crawl can replicate. Dense, edited, authoritative." The outrage hit a fever pitch this week but news of the practice broke last January, when court documents exposed "Project Panama," a program within Anthropic to scan then destroy as many books as they can get a hold of. Anthropic settled with authors for $1.5 billion, but the unsealed filings showed that the practice is considered legal, just bad publicity, and the company was willing to break the bank to keep it under wraps. Now the juice is out of the tube, and suspicious rare book orders are under intense scrutiny. "It benefits me financially as well as by clearing out old inventory that is otherwise unlikely to sell," one anonymous seller told 404's Samantha Cole. "On the other hand, I don't like the end-use, and I don't like that uncommon books are being pulped." Confronting the hell they've rendered for themselves, AI firms are struggling to find "clean" source material to train LLMs on. Whether it's low quality web content, or material already generated by AI that threatens negative feedback loops, these companies are trying to siphon higher quality stuff without raising too many alarms. Earlier this summer A24 announced a partnership with Google to assist training DeepMind, hoping to help their media generators escape the bog of muddy brown, cursed Ghibli slop. As for the destruction part, it's unlikely being done in attempts to cover their tracks or some Brianiac-style rare knowledge obsession. The Project Panama documents didn't specify why they trashed the books, but it's most likely just trying to save a buck. Archiving and scanning books doesn't have to destroy the source material, but being done cheaply and quickly, it likely will. Ripping apart the spines for clearer scans and disposing of the crumpled heap.
[13]
Buy, scan, destroy: AI firms are shredding millions of books to train their chatbots
AI companies are buying, cutting apart and scanning millions of physical books to train chatbots, according to a Washington Post report based on newly unsealed court filings. The documents reveal Anthropic's secret "Project Panama" and shed light on similar book-acquisition efforts by Meta, OpenAI and Google, as the AI industry's race for high-quality training data fuels an escalating copyright battle. Artificial intelligence companies are buying millions of physical books, cutting them apart, scanning every page and recycling the remains to train the chatbots powering today's AI boom, according to a report by The Washington Post, based on newly unsealed court filings in a copyright lawsuit against AI startup Anthropic. The filings also reveal how the race to build more capable AI models pushed companies to secure vast collections of books, setting off a wave of copyright lawsuits. The court documents provide one of the clearest glimpses yet into the AI industry's search for high-quality training data. While Anthropic's internal book-scanning project is at the centre of the filings, separate lawsuits involving Meta, OpenAI and Google also highlight how leading AI companies sought access to millions of books to improve their models, although the methods differed across companies. From bookshelf to chatbotAt the centre of the revelations is Anthropic's internal initiative called "Project Panama", which company documents described as an effort to "destructively scan all the books in the world." According to the report, the company spent tens of millions of dollars buying millions of used books, slicing off their spines, scanning every page and sending the remains for recycling to create training data for its Claude AI models. The planning documents also stated: "We don't want it to be known that we are working on this." According to the report, Anthropic initially explored sourcing books from libraries and used bookstores, including New York's Strand Book Store, before eventually purchasing large batches from used-book retailers such as Better World Books and the UK's World of Books. A proposal from one of its scanning vendors said the company planned to digitise between 500,000 and 2 million books in six months, using industrial cutting machines to remove bindings before scanning the pages and sending the remains for recycling. To lead the effort, Anthropic hired Tom Turvey, a former Google executive who helped create Google's Google Books project more than two decades ago, according to The Post. The AI arms raceBooks were viewed as a critical resource because they offered higher-quality writing than much of the internet. One Anthropic co-founder wrote internally that books could teach AI models "how to write well" instead of imitating "low quality internet speak," according to the report. Before launching Project Panama, Anthropic employees had also downloaded books from shadow libraries such as LibGen and Pirate Library Mirror, which host copyrighted material without permission. Anthropic has said it never trained a commercial AI model using the LibGen dataset and did not use Pirate Library Mirror to train any complete AI model. The filings suggest Anthropic was not alone. According to The Post, Meta employees discussed using LibGen to obtain millions of books, with internal messages showing concerns over copyright risks. One engineer reportedly wrote, "Torrenting from a corporate laptop doesn't feel right," while another discussion referred to using rented servers to avoid the activity being traced back to the company. Meta has denied illegally distributing copyrighted works. OpenAI has acknowledged downloading LibGen but told a court it deleted the files before the release of ChatGPT. Google is also facing copyright litigation related to AI training data, with two major publishers recently seeking to join an existing lawsuit against the company. The copyright plot twistThe legal battle over AI training data is far from settled. In June, a US judge ruled that using books to train AI models could qualify as fair use because the process is "transformative." However, the judge also found Anthropic could still face liability over how it acquired some of the books by downloading pirated copies. Anthropic later agreed to pay $1.5 billion to settle claims related to the acquisition of books, while maintaining that the settlement concerned the method of acquisition rather than the legality of AI training itself. Most copyright cases involving AI companies, authors and publishers are still working their way through US courts, meaning the broader legal boundaries for training AI on copyrighted material remain unresolved.
[14]
AI burns books? Tech giants are buying, shredding books to train smarter chatbots
AI companies are buying and then destroying millions of rare books to prevent AI slop content that consumers are vocally against, 404 Media reported last week. Silicon Valley tech giants are paying companies and contractors to buy up rare books, which are then scanned in a high-speed machine that cuts their spines out and then shreds the originals. The tech giants are reportedly buying up the rare books to train new AI models and prevent AI "slop." In one article on its site, ISBNdb, a company that claims to have the "the world's largest book database," argued that books published before 2022 were best for AI training data because they would not have any AI-generated text. In a report uncovered by the Washington Post in January, one Anthropic co-founder suggested that feeding AI models books could teach them "how to write well" instead of producing "low-quality internet speak." "The world's best AI training data is sitting on a shelf," ISBNdb wrote in a since-deleted blog post. "Print books from the pre-LLM era are structurally guaranteed to be free of this contamination. That alone is a significant advantage [...] "Physical books published before this date [pre-2022] are structurally clean of modern poisoning tools." 404 Media reported that the AI tech companies were interested in buying up books to prevent "model collapse," where AI models trained on lower-quality AI-generated results progressively lose quality. Executives believed that vast troves of books were essential to give the models new data and prevent their model collapse. Why are AI companies destroying old books? Notably, special edition book sellers say that they usually sell one or two books to a single customer. "In the rare book trade, it's very seldom that people want to buy more than one book," antique book seller Pieter de Vries told The Telegraph in an interview published last week. "So if somebody comes and says, 'I want a couple of hundred of your books,' it's very strange." But ISBNdb and companies like it are now helping AI tech giants purchase orders ranging from 1,000 to one million books. The AI companies' attempt to hoover up books, art, news articles, and other forms of media has gotten attention in several instances since 2024. Most notably, the Washington Post reported in January on Anthropic's attempts to buy millions of books, slice their spines, and scan their pages to feed more data into the company's chatbot, Claude. Anthropic's legal battle over AI and book copyrights That instance later led to a multi-million dollar class action lawsuit in which several authors sued Anthropic, arguing that the company, which is backed by Amazon and Alphabet (Google's parent company), used pirated versions of their books without permission to teach Claude to respond to human prompts. Anthropic settled with the authors at $1.5 billion. Notably, though, the fact that Anthropic destroyed the books made the company's case stronger. Judge William Alsup ruled last June that Anthropic made fair use of the authors' work to train Claude, but found that the company violated their rights by saving more than seven million pirated books to a "central library" that would not necessarily be used for AI training. "Here, every purchased print copy was copied in order to save storage space and to enable searchability as a digital copy," Alsup wrote. "The print original was destroyed. One replaced the other. And, there is no evidence that the new, digital copy was shown, shared, or sold outside the company," he said. In short, because the AI company pulped the books, the digital version that existed afterward "replaced" the physical one. The judge ruled that the companies merely transformed the books, which meant that it was not a violation of US copyright law. The tech giants clearly wanted this secret since the beginning, because the optics of destroying books are unpalatable for many. Documents uncovered by the Washington Post found that even internally, companies wanted to distance themselves from the effort. "Project Panama is our effort to destructively scan all the books in the world," one internal document unsealed in legal filings last said, as reported by the Washington Post. "We don't want it to be known that we are working on this." As ISBNdb wrote in a since-deleted blog post on its website: "The optics problem is real. 'AI company destroys two million books' is not a headline that generates sympathy." Tech giants, former officials decry AI companies for destroying books But the headlines and lawsuits became public, and the public is rather unsympathetic. "So AI labs are buying old books by the pallet, slicing them apart, scanning the pages, and pulping what's left," former US House rep. Brad Carson wrote in a post on X/Twitter. "Here's the perverse part. A federal court blessed this precisely because the original is destroyed. One legal copy replaces another, so it's fair use. Whatever you think of that ruling or fair use, notice what it does. The law now rewards destruction and penalizes preservation. "A lab that wants to scan a book and keep it, or donate it, or deposit the scan in a public archive, has weaker legal footing than a lab that shreds everything. We have built a legal machine that pays people to pulp books and punishes them for saving them." Some tech giants have voiced their distaste for the destruction of the books. "I've asked the SpaceX AI team to preserve any rare books in a library and scan them the hard way rather than just cutting off the spine and scanning," Elon Musk said in an X/Twitter post. "There's something particularly misanthropic about the mechanized destruction of such intimate human objects," said SEO of Factory AI Matan Grinberg. "History seldom looks kindly on those who destroy books, whatever the reasons." One critic of the AI industry's approach to copyrighted work and the founder of Fairly Trained, a creator-rights group, Ed Newton-Rex, told the Telegraph that the secrecy with which companies like ISBNdb and Anthropic acquire books is damning. "Clearly both the provider of these books and the AI companies know that this is a terrible look and they don't want the specifics to get out," he told the Telegraph. "If you are just going and spending $1 on a used book, with all of the money going to a book wholesaler, should that give you the right to train a commercial generative AI model on that book, which will then be able to compete with the author who wrote it?" Newton-Rex said. "A lot of people, myself included, think it shouldn't." "There is surely no more fitting image in the generative AI age for the exploitation that underlies this technology than almost trillion-dollar companies buying books for a few cents or a dollar each, scanning them, training on them, then destroying them, essentially subsuming culture," he added. In response to the flurry of reporting around the destroyed books, ISBNdb disputed the reports that it helped purchase large volumes of books for AI training. "We've seen the recent coverage about a marketing landing page on our site, and we understand the concern it raised," the company wrote in a statement. "The facts: ISBNdb has never purchased, scanned, or sold a book - for AI training or anything else. We don't train AI models, and we never have. The page was a test of market interest; no such service was ever brought to life. We've taken the page down. "Our job is helping people find books. For more than two decades, ISBNdb has been the card catalog of the book world - the data behind how bookstores, libraries, and reading apps connect readers with titles. Data about books, not the books themselves. That hasn't changed."
[15]
AI's hunger for data is now consuming rare books and booksellers are cashing in
Artificial intelligence firms are purchasing vast quantities of physical books. These books are scanned to train large language models and then destroyed. Rare and out-of-print titles are particularly sought after for this purpose. Booksellers report a significant surge in demand for these older publications. This practice raises concerns about the loss of scarce literary works. Artificial intelligence companies are increasingly turning to physical books, including rare and antique editions, to feed their large language models, creating a booming market for bulk book purchases that is raising concerns about the destruction of scarce literary works, according to a report by 404 Media. The report says AI companies are buying thousands, and in some cases millions, of used books, scanning their contents into digital datasets and then destroying the physical copies. The practice has gathered momentum as developers seek cleaner, high-quality training material that predates the explosion of AI-generated content online. Also read: Destroyed in 2005, reborn through LNG: Inside America's unlikely energy powerhouse The trend follows legal developments in the United States that have allowed companies to digitise books they legally purchased under the first-sale doctrine. In a lawsuit that has since been settled, court documents showed that Anthropic used hydraulic cutting machines to remove pages from books before scanning them with industrial imaging equipment for AI training. A judge found the digitisation process to be "transformative," making it eligible for fair-use protection because the books were converted into digital form rather than redistributed as new copies. Rare and out-of-print books fuel AI data raceAccording to 404 Media, businesses that once primarily served libraries and booksellers are now actively courting AI firms. One such company, ISBNdb, which describes itself as operating the "world's largest book database," argues that older printed books offer an advantage because they were published before AI-generated text became widespread. "Print books from the pre-LLM era are structurally guaranteed to be free of this contamination. That alone is a significant advantage," the company says on its website, as quoted by 404 Media. It further states: "Physical books published before this date [pre-2022] are structurally clean of modern poisoning tools." According to the report, ISBNdb now facilitates bulk purchases ranging from 1,000 to one million books per order for AI companies. The company also advertises anonymous procurement services, acknowledging the reputational risks surrounding the practice. "The optics problem is real," ISBNdb says on its website. "'AI company destroys two million books' is not a headline that generates sympathy." Also read: In 2006, Canada launched a mission to save Atlantic salmon. Nearly 20 years later, the fish are returning Booksellers see windfall but worry about what's being lostThe surge in demand is already being felt by independent booksellers. One seller told 404 Media that weekly sales jumped from fewer than 20 books to hundreds in April, with buyers selecting seemingly random titles that all carried ISBN numbers, leading the seller to believe AI companies were behind the purchases. The bookseller said many of the works being sold were rare or out of print, raising fears that some of the last remaining physical copies could be disappearing. "I personally have mixed feelings about all of this," the bookseller told 404 Media. "It benefits me financially as well as by clearing out old inventory that is otherwise unlikely to sell. I've been well-suited for these sales with inventory from overseas and foreign language books. On the other hand, I don't like the end-use, and I don't like that uncommon books are being pulped." According to the report, similar buying patterns have been reported by rare-book dealers in the Netherlands, where sellers say they are receiving unusually large orders that they suspect originate from AI companies. While booksellers often cannot verify the identity of buyers because intermediaries such as ISBNdb keep purchasers anonymous, many in the industry believe the bulk acquisitions reflect the AI sector's growing appetite for high-quality, pre-AI training data.
[16]
Project Panama: Scanning And Book Destruction At Anthropic
Book burning. Book pulping. Book vandalising. It's all the fashion and, dare one say it, the rage. Libraries are carting them off to the dump. Repositories of memory are being shredded in favour of supposedly more useful digital formats. And now, Anthropic's hungering for books, not as sources of enlightened knowledge so much as blue raw data for their Language Learning Models, is there for all to see. Acquire the books in question. Give service providers the task of severing their spines. Employ scanners to process the information. Dispatch the paper to be pulped and recycled (awfully good of them). The result: a private research library able to nourish Claude, the company's premier LLM. Last month, an investigation by 404 Media took note of the purchasing habits of AI companies as regards old books - whatever that means - identifying ISBNdb as an instrumental broker in the field. The company, according to its own description, "gathers data from various public sources like libraries and merchants to compile a vast collection of unique book data searchable by ISBN, title, author or publisher." At present, it boasts 111,978,817 searchable books and offers clients somewhere between 1,000 to 1 million books per engagement. Older print publications are advertised as the purer sort, uncontaminated by presence of generative-AI. "Print books from the pre-LLM era are structurally guaranteed to be free of this contamination," states an article published by the company. It remains unclear whether ISBNdb's client list is crowded by those of the destructive scanning persuasion, though the company is not oblivious to the problem. The same goes for other book sellers as to whether their activities are feeding the AI maw of physical pulverisation. The Canadian company Zoom Books specialises in acquiring nonfiction and academic titles from the 1970s in bulk, storing titles in European warehouses (Germany, in particular), before shipping them off to the US and Canada for destructive scanning. The practice struck Thomas Koch, press spokesperson for the German Publishers and Booksellers Association, as particularly reprehensible, though he admitted a note of caution about the extent this was taking place : "It appears to be yet another example of AI companies using vast quantities of copyright-protected works to train their language models, without consent and without payment." Anthropic's activities came to light in filed documents in the northern California District Court case of Bartz v Anthropic PBC, involving the authors Andrea Bartz, Charles Graebner and Kirk Wallace Johnson. All alleged copyright infringements against the company. Known as "Project Panama", this extract-and-pulp mission was led by Tom Turvey, a previous employee of Google who had made strides in creating such partnerships as Google Books. As the introduction to the court order notes, the project was intended to create "a central library of 'all the books in the world' to retain 'forever.'" The library would enable the AI firm to select "various sets and subsets of digitized books to train various large language models under development to power its AI services." The company's appetite for book purchases proved voracious. Instead of brokering agreements with publishers to license copies for purposes of training AI (this approach, derided as "legal/practice/business slog", was initially explored, if only tentatively), "Turvey and his team emailed major book distributors and retailers about bulk purchasing their print copies for the AI firm's 'research library'." Millions of dollars were expended on printed books, even those in used condition. Retained service providers then went about their work: stripping the book bindings, cutting the pages to size, scanning the books into digital form. Paper originals were discarded. "Each print copy book resulted in a PDF copy containing images of the scanned pages with machine-readable text (including front and back cover scans for softcover books)." The company also made a point of harnessing pirated digital books, which eventually led the company to reach a $1.5 billion out-of-court settlement. These included five million copies from LibGen, two million items from Pirate Library Mirror and approximately 183,000 from Books3. The court describes how Anthropic "came to value most highly for its data mixes books like the ones" written by the plaintiff authors "because of the creative expressions they contained." Those using the Claude LLM wanted it "to write as accurately and compellingly" as the authors, using "well-curated facts, well-organized analyses, and captivating fictional narratives". On April 13, 2024, an internal memorandum was circulated in the company which let the mad cat of destruction out of the bag. "Project Panama is our effort to destructively scan all of the books in the world," is the lavishly ambitious claim heading the message by Turvey. The memorandum in question urges discretion. "Why use a codename? We use 'soft codename' because we don't want to be known that we are working on this. This document is available to all Anthropic employees, but you should avoid talking about it in public areas and the fact that we are working on this should not be shared with anyone outside Anthropic." In an email to Snopes, a spokesperson for Anthropic issued something of a qualification, claiming that, "None of our data acquisition programs buy and destroy 'rare' or 'antiquarian' books." The company has made a distinction between "rare and valuable", a formulation most shoddy. In all this excitement of vandalism, it was easy to forget that the plaintiffs had failed to convince Judge William Alsup that copyright violations had taken place. "Anthropic's LLMs have not reproduced to the public a given work's creative elements, nor even one author's identifiable expressive style". Claude "outputted grammar, composition, and style that the underlying LLM distilled from thousands of works." But copyright did not cover "'method[s] of operation, concept[s], [or] principle[s]' 'illustrated' [ ] or embodied in [a] work." The copyrighted works, in being used to train LLMs in generating "new text was quintessentially transformative." On the issue of the assembled "research library", the court found no issue with Anthropic's conversion of each book copy's format from print form to digital, effectively turning a blind eye to the eventual destruction of their physical existence. "Anthropic purchased its print copies fair and square." But the pirated copies were something else. "The person who copies the textbook from a pirate site has infringed already, full stop." Anthropic's claim that using copies for a central library was a case of fair use was untenable. In all, a melancholic day for books, an undeserved if partial victory for Anthropic, and the offering of a bleak, wearying prospect that the creative faculties of the writer risk being turned into the blind and dumb receptiveness of the slothful consumer. The means of language production, as Kathryn James, rare book librarian of Yale University's Lillian Goldman Law Library at Yale University appositely remarks, has been ceded. Dr. Binoy Kampmark was a Commonwealth Scholar at Selwyn College, Cambridge. He currently lectures at RMIT University. Email: [email protected]
Share
Copy Link
AI firms are buying and destroying millions of physical books to train large language models like Claude AI. Anthropic's secretive Project Panama uses destructive scanning—cutting spines and pulping originals—to digitize pre-2022 texts free of AI-generated content. Booksellers worldwide report mysterious bulk orders, raising fears that irreplaceable rare volumes are being lost forever.
Court documents from the copyright case against Anthropic have revealed a secretive operation called Project Panama, described in internal memos as "our effort to destructively scan all the books in the world."
1
4
The AI company explicitly advised employees to use a codename "because we don't want it to be known that we are working on this."4
Documents unsealed in the Bartz v Anthropic PBC case show that in spring 2024, Anthropic spent millions of dollars purchasing millions of new and used books specifically to destroy them for AI training data.1
The destructive scanning process involves stripping books from their bindings, cutting spines and edges, feeding individual pages through high-speed industrial scanners, and discarding the paper originals as trash.
3
4
Anthropic hired logistics managers, sourced books from vendors, and housed them in warehouses where staff moved between neatly stacked and labeled shelves before the volumes were shredded and disposed of.4
The company settled with authors for $1.5 billion after a class action lawsuit revealed it had kept 7 million pirated books in a central library.3

Source: Tom's Guide
AI companies are specifically hunting for books published before 2022 because these texts are guaranteed to be free of AI-generated content.
1
5
Large language models like Claude AI need "well-curated facts, well-organized analyses, and captivating fictional narratives" to improve their ability to predict language outcomes and write as compellingly as human authors.4
Training AI on its own output can cause model collapse, a phenomenon where models degrade significantly in quality when fed machine-generated text.1
Books represent one of the best sources for complex, high-quality, long-form text that is "dense, edited, authoritative," unlike internet content increasingly filled with AI slop.
5
Pre-2022 volumes also avoid deliberate data-poisoning techniques that authors increasingly use to disrupt AI training.5
Anthropic chose destructive scanning over securing copyright permission for existing e-books because dealing with copyright represented a "legal/practice/business slog" that the company preferred to avoid.4
Booksellers worldwide have reported strange waves of unusually large orders through platforms like AbeBooks and Biblio.
2
Charlie Becker, manager of Becker's Books in Houston, told reporters that 95 of the past 100 books ordered through one platform went to a single buyer, with one order totaling 70 books.1
The orders are price-insensitive and oddly random—one Australian bookseller received three boxes in May pairing a 1970s soil-mechanics manual with a $9 poetry book that had sat unsold for 20 years.2
The books are almost entirely nonfiction, mostly published from the 1970s to the 1990s, written in various languages, and include titles like "The Insider's Guide to Metro Denver from 1995" or "How to Use Corel WordPerfect 1991."
1
In many cases, shipping costs exceed the books' resale value.1
Australian secondhand booksellers traced many orders to Zoom Books, a Canadian "book recycler" that offers strict non-disclosure agreements keeping customer identities confidential.2
5

Source: Fortune
While most books in bulk orders appear to be common out-of-date titles, booksellers fear that irreplaceable rare volumes are entering the pipeline.
2
Delfina Manor in Victoria drew the line when she realized some orders might include only known copies of certain works.2
Nick Dawes, who keeps about 500,000 books, said he would happily sell in bulk but refuses if "this is the only known copy, I'm going to cut it up."2
ISBNdb, a commercial book metadata service, was named in 404 Media reports as potentially supplying up to 1 million books per order to AI labs.
1
5
However, no bookseller could confirm orders from ISBNdb, and the company has since scrubbed its website of AI training references, denying it ever bought, scanned, or sold books for this purpose.1
3
Zoom Books says it resells books intact and only recycles what cannot be rehomed, but won't reveal whether it sells to AI firms, citing commercial confidentiality.2
Related Stories
In July, Judge William Alsup ruled that training AI on books constitutes fair use under copyright law because the practice is "transformative."
1
3
The court found that destroying each printed copy during scanning meant "one replaced the other" rather than creating multiple copies, making it legally permissible.4
5
The ruling noted there was "no evidence that the new, digital copy was shown, shared, or sold outside the company."4
This legal precedent means destructive scanning is now both lawful and less regulated than strip mining in the United States, which has very few legal provisions governing the treatment or export of books as cultural heritage.
4
Anthropic spokesperson told Tom's Guide that "sourcing books is a widely used approach for training large language models across the AI industry" and that the company buys "books from regular commercial markets" rather than rare or antiquarian titles.3
However, ethical concerns persist about cultural preservation when opacity prevents booksellers from screening buyers.2

Source: TechRadar
Michael Burry called the practice "evil incarnate," while Elon Musk responded on X by saying he asked engineers training xAI's models to "preserve any rare books in a library and scan them the hard way."
1
5
Critics note that unlike websites or modern publications, exceptionally scarce historical works that survived wars, fires, and centuries of handling cannot be reproduced after physical copies disappear.5
The scale of destruction has been compared unfavorably to historical losses: "Such a level of destruction is an order of magnitude bigger than the loss of the Library of Alexandria. Yet, it is unfolding with none of the outrage that history reserves for burned libraries."
5
With the court ruling establishing precedent and AI companies across the industry reportedly adopting similar practices, watch for increased tensions between the tech sector's hunger for clean training data and cultural preservation advocates demanding transparency in book sourcing.Summarized by
Navi
[1]
[2]
[3]
[4]
23 Jul 2026•Policy and Regulation
26 Jun 2025•Technology

12 Aug 2026•Technology

1
Technology

2
Technology

3
Policy and Regulation
