5 Sources
[1]
AI companies are buying and destroying old books for training data
AI companies have spent years pulling training material from the internet. Now, as more of the web fills with AI-generated writing, some are looking for text in a place largely untouched by chatbots: used-book shelves. Until July 28, ISBNdb, a company that operates an online book database, advertised its ability to source as many as 1 million physical books per order for AI developers. The since-deleted page offered books "tailored to your LLM training needs," according to 404 Media, including older, specialized, rare, and out-of-print titles gathered from used bookstores and other catalogs. The company also promised confidentiality. Its marketing materials advertised a "strict NDA on every engagement" and said clients' identities, strategies, and acquisition targets would not be disclosed. That secrecy attracted attention because converting physical books into AI training data can involve destructive scanning: cutting off their bindings, feeding the loose pages through industrial scanners, and discarding or recycling the originals. By July 28, ISBNdb had removed the sourcing page and its promise of an NDA. The company said the service had been part of "exploring demand" and that it had "chosen to pivot away from that direction," in a news update. Still, ISBNdb's brief sales pitch offered a glimpse into a book-buying operation that one major AI developer has already carried out on a much larger scale. Federal court records show that Anthropic purchased and scanned millions of physical books, while used-book sellers in the United States and Europe have recently reported unusual bulk orders for obscure and out-of-print titles. The reports have opened several practical questions: Why have older books become so valuable to AI developers, how widespread is destructive scanning, and what happens to the originals once their pages become training data? Why AI companies are hunting for old books Older books offer something that has become surprisingly difficult to guarantee online: writing produced entirely by humans. Large language models are trained on enormous quantities of text collected from websites, articles, books, code repositories, and other digital sources. Since generative AI tools became widely available, however, the internet has filled with machine-written summaries, product listings, social posts, and articles. That creates a problem for developers assembling new training datasets. If a model is trained too heavily on material produced by earlier models, it can begin reinforcing their mistakes while losing some of the variety and less common information contained in the original human data. Researchers call the process "model collapse." A physical book printed before the current AI boom offers a relatively clean alternative. Unlike a webpage that may have been quietly generated or rewritten by a chatbot, an older book provides a more dependable record of human writing. Books are also edited, structured, and often contain specialized information that cannot easily be found elsewhere online. ISBNdb leaned heavily on those qualities in its sales pitch. "Books represent curated, peer-reviewed, domain-specific human knowledge, structured in a way no web crawl can replicate," the company wrote. "Dense, edited, authoritative." ISBNdb also emphasized that many potentially useful titles had never been fully digitized. Its marketing effectively presented the pre-chatbot bookshelf as a reserve of human-created training material that had yet to be mixed with AI output. That demand did not appear out of nowhere. It follows a much longer effort to turn printed books into searchable digital information. AI didn't invent mass book scanning People have been converting books into digital files for decades. Project Gutenberg began creating electronic versions of public-domain works in 1971 and now offers more than 75,000 free ebooks. The industrial-scale version arrived with Google Books in 2004. Working with libraries and publishers, Google scanned books and made their contents searchable online. By 2019, the company said it had assembled a collection of more than 40 million books in over 400 languages. Those scans also helped form HathiTrust, a digital repository created by research libraries in 2008 to preserve their collections and make as much of the material publicly accessible as copyright law allowed. The project prompted an earlier version of today's copyright fight. Authors sued Google for scanning copyrighted books without permission, but a federal appeals court ruled in 2015 that the project qualified as fair use. The court found that creating a searchable index and displaying limited snippets gave the books a new purpose without providing readers with a replacement for the originals. Other digitization programs have emphasized preservation and access. The Internet Archive, for example, says it has digitized more than 25 million books since 2006 using nondestructive scanning designed to keep the bound volumes intact. But converting a book into data does not necessarily make it publicly accessible. That distinction concerned Charlie D. Becker, a second-generation bookseller whose family runs Becker's Books in Houston. His store recently received a single order for 70 obscure titles, including a 1995 guide to metropolitan Denver and manuals explaining how to use WordPerfect in 1991. After investigating, Becker suspected the buyer was using an algorithm to find underpriced books and relist them on Amazon, rather than acquiring them for AI training. Most of the titles were either unavailable on Amazon or listed there for as much as 20 times his store's price, and the shipments appeared to be going to Fulfillment by Amazon preparation companies. Still, Becker said an obscure book can be especially easy to lose because few people consider it worth preserving. Books that fail to sell through Amazon's fulfillment system may eventually be liquidated, while some titles have little more than scattered listings across private databases to prove they existed. "Everyone assumes the internet preserved everything," Becker wrote in a longer account of the orders. "It didn't." Anthropic has already scanned millions of books The buyers behind the recent bookstore orders remain unclear. The destructive scanning process, however, is no longer hypothetical. In 2024, Anthropic hired Tom Turvey, a former Google executive who had worked on partnerships for the Google Books project. According to a June 2025 federal court ruling, Turvey was tasked with helping the company obtain "all the books in the world" for an internal research library. Turvey initially contacted publishers about licensing their books, but those conversations did not continue. His team then approached major book distributors and retailers about purchasing print copies in bulk. Anthropic ultimately spent millions of dollars acquiring millions of physical books, many of them used. Service providers removed the books from their bindings, cut their pages to the appropriate size, and fed them through scanners. The searchable PDF files went into Anthropic's internal library. The paper originals were discarded. Engineers could then select groups of those books for inclusion in datasets used to train the large language models behind Claude. That process became central to a legal battle over how Anthropic obtained its training material. In June 2025, U.S. District Judge William Alsup ruled that the company's use of books to train Claude was transformative and qualified as fair use under the specific circumstances of the case. Alsup also found that converting legally purchased print books into digital files could qualify as fair use because Anthropic destroyed each physical copy and replaced it with one internal digital copy. In other words, the company did not keep both versions. The ruling did not give AI companies blanket permission to copy any book they could find. Alsup drew a sharp distinction between the print books Anthropic had legally purchased and the more than 7 million pirated books the company had downloaded and stored in a permanent digital library. Purchasing physical copies of some titles later did not erase the original piracy, he found. That distinction eventually became expensive. On July 20, a federal judge approved Anthropic's $1.5 billion settlement with authors and publishers, resolving claims involving approximately 482,000 pirated books. Eligible rights holders are expected to receive about $3,000 per title, according to Reuters. For some authors, the payment does not resolve the larger disagreement. Charles Graeber, one of the case's original plaintiffs, told NPR that he was proud authors had secured a substantial settlement, but said the case had cost him more than two years of time, travel, and professional opportunities. Fellow plaintiff Andrea Bartz questioned a system that allows companies to train commercial models on legally purchased books without negotiating separate licenses with the people who wrote them. "The algorithm is being used to essentially try to put us out of a job," she told NPR. Anthropic has maintained that training AI models on books is protected by fair use. The court's ruling nevertheless helps explain why physical books may be especially attractive to developers: A lawfully purchased copy gives the company a much stronger legal position than a file downloaded from a pirate library. Destroying the original may be a legally useful distinction. For many readers, it is also the most unsettling part of the story. Who is buying all these books? Anthropic's operation is documented in court records. The source of the more recent bookstore orders is harder to pin down. One bookseller specializing in uncommon and low-circulation titles told 404 Media that his weekly sales jumped from roughly 20 books during a good week to several hundred after the orders began arriving in April. The requested titles did not appear to share a subject, author, genre, or language. They did, however, all have International Standard Book Numbers, or ISBNs, leading the seller to suspect that they had been selected through a book database. The surge was financially helpful and allowed him to clear inventory that might otherwise have remained unsold. He was less enthusiastic about where the books might be going. "I don't like the end-use, and I don't like that uncommon books are being pulped," he said. Booksellers in Europe have reported similarly broad requests. An antiquarian bookseller in the Netherlands received a list of approximately 3,000 English-language books organized by ISBN, ranging from an academic study of Irish folklore to a technical book about laser shock peening. There is no public confirmation that every unusual bulk order came from an AI company or that every book purchased through these orders was destroyed. There is also no evidence that developers are intentionally hunting for the final surviving copies of rare titles. The uncertainty itself is part of the concern. A company purchasing from a massive ISBN list could sweep up uncommon or out-of-print editions without first checking how many physical copies remain. Why the story struck a nerve Once the reports reached social media, they were accompanied by an unsettling visual: a cutting blade moving inch by inch through a book's spine. Comparisons to Fahrenheit 451 and the Library of Alexandria followed quickly. Book destruction carries a particular weight because it has historically represented more than the loss of paper. The 1933 Nazi book burnings targeted works deemed "un-German," including books by Jewish, pacifist, and left-wing writers. The act became an enduring symbol of censorship and the suppression of ideas, according to the United States Holocaust Memorial Museum. That symbolism is central to Fahrenheit 451, Ray Bradbury's 1953 novel about a society where books are outlawed and burned. Bradbury said his warning extended beyond government censorship to television reducing knowledge to digestible fragments and eroding interest in reading. Users have also invoked the Library of Alexandria, another enduring symbol of lost knowledge, although historians believe it declined gradually through political upheaval, reduced support, neglect, and repeated damage rather than disappearing in one catastrophic fire. Elon Musk joined the conversation on July 27. He wrote on X that he had asked the SpaceXAI team to preserve rare books in a library and scan them "the hard way," without removing their spines. As the discussion spread, users resurfaced a 2011 Cracked article about libraries, universities, and retailers destroying unwanted books years before the current AI boom. That history has informed a less alarmed response. Some users argued that the books shown in warehouses appeared to be ordinary, unwanted inventory rather than irreplaceable artifacts. If a book would otherwise be recycled without being read again, they asked, could scanning it first preserve something that would have been lost? AI adds a complication. The text may survive, but inside a private company's research library rather than a public archive. A forgotten manual or travel guide may have little resale value, yet still contain exactly what an AI developer wants: edited human writing created before the flood of chatbot output. That has led authors and publishers to argue that they should have more control over how their work is used, especially when it is helping companies build commercial products. AI developers, meanwhile, continue to argue that training a model is a transformative use of the material rather than a replacement for the original books. ISBNdb has removed its sourcing page, but the demand behind it remains. The web's AI problem has sent developers back to the bookshelf. The next chapter will depend on whether they can extract what they need without leaving those shelves any emptier.
[2]
This Dutch bookseller thought a request for 3,000 copies was 'spam or phishing.' Instead, AI companies are scanning and destroying books to train AI | Fortune
It was just another sunny summer day in Haarlem, Netherlands, when an unusual request landed in Dutch antiquarian bookseller Pieter de Vries' inbox. A woman named Natalia, identifying herself as being with a company called "2077AI," was looking to source what she described as a "fairly large order" of books. Attached was a spreadsheet listing more than 3,000 ISBNs along with instructions to match titles, prepare quotes and estimate shipping costs directed to China. De Vries barely looked at it, dismissing the message as spam or phishing without reading past the first few lines. It wasn't until weeks later, when a Dutch journalist investigating the procurement request contacted him, that he learned of the request's connection to artificial intelligence. "I was shocked!" said de Vries, who provided Fortune with the email and subsequent spreadsheet containing 3,001 titles, mostly published between 2020 and 2021 by academic publishers like Emerald Publishing, Elsevier, Wiley, Routledge and Oxford University Press, and ranging from business and education to engineering, public policy and medicine. "Very unusual," de Vries said. "That's why I regarded the request as spam or phishing." De Vries wasn't the only one. He told Fortune several other Dutch antiquarian booksellers received the same request and, like him, dismissed it. (They didn't take "it seriously" and "considered it as scam," he said). While the inquiry had little effect on his business, he said, it could have been a different story for booksellers who specialize in selling secondhand books at scale. "I deal in old and rare books," de Vries said, adding he rarely receives book orders by e-mail and if he does it's at most for only one book. "These books are expensive and exclusive. So I do not do bulk trade." AI's appetite for books The email De Vries received never mentioned artificial intelligence or explained what the books would be used for., but it surfaced as AI companies increasingly look beyond the open internet for high-quality written material to train large language models. Last summer, public attention to AI companies' use of books intensified when court records revealed Anthropic had purchased millions of physical books, removed their bindings, scanned them and discarded the originals to build a searchable digital library used to train its AI models in what internal documents called "Project Panama," according to The Washington Post. The process, known as "destructive scanning," involves cutting the spine from a book so its pages can be fed through high-speed scanners before the remaining physical copy is discarded. The case later settled after a federal judge ruled that using legally purchased books to train AI models constituted fair use under copyright law. Separate claims involving Anthropic's downloading of books from the LibGen and PiLiMi online libraries were resolved through the settlement. According to an Anthropic spokesperson, Claude models are trained on a mix of publicly available web data, commercially acquired datasets and internally generated data. Anthropic purchases books through regular commercial markets, and none of its data acquisition programs acquire or destroy rare or antiquarian books. The hidden supply chain The court records showed what happened after books were acquired, not how companies source millions of physical volumes in the first place. Earlier this year, 404 Media reported that ISBNdb, a company best known for maintaining metadata tied to International Standard Book Numbers (ISBNs), the unique numerical identifiers encoded in the barcodes found on most books, had begun advertising services to help AI labs source bulk printed book purchases ranging from 1,000 to 1 million books. According to the report, the company's website said the service was tailored to "LLM training needs" and "delivered at the scale AI demands." The webpages have since been removed. ISBNdb later told Fortune the service was never launched, it never purchased or scanned books for AI companies and the webpages reflected an exploratory concept rather than an active business. Similar reports have emerged in Germany and Switzerland, where local media reported that secondhand booksellers received unusual bulk orders for highly specialized books that appeared inconsistent with traditional collecting or resale. Zoom Books, a Canadian company identified in some of the reports, denied these allegations and told Swiss broadcaster SRF the purchases were part of its regular recycling and trading model. For an industry built on algorithms, the email offers a rare glimpse into something much more tangible: the physical supply chain of books that may ultimately power AI models. Fortune contacted 2077AI seeking comment about the purpose of the procurement request and whether the books were intended for AI training. The company did not respond to requests for comment before publication.
[3]
AI companies are reportedly buying up and destroying millions of old books for training
According to a new report from 404 Media, AI companies are not only using books to build and train powerful AI models, but they're also destroying them. Yes, to digitize millions of books in a way that saves them time and money, AI companies are removing spines and individual pages to make the whole process faster. It's a system referred to as 'destructive scanning,' where the source material is effectively destroyed. This process came to light a year ago, when Anthropic's "Project Panama" initiative was taken to court in the United States for this very reason. Even though the court ruled in favor of Anthropic because it legally purchased these books to turn them into digital copies, it led to widespread backlash because it meant that millions of books were destroyed, everything from paperbacks to rare prints and editions. Even though destructive scanning is viewed as 'fair use' by federal courts, AI companies are now looking to shield themselves from criticism by working with intermediaries to acquire books at scale on their behalf. The 404 Media report notes that these intermediaries promise confidentiality as they source physical copies of books in vast quantities, in the range of hundreds of thousands, or more, at a time. One of these companies is ISBNdb, which boasted that it could source rare books from "library shelves, used bookstores, and out-of-print catalogs," with the option to buy 1 million books at a time with a "strict NDA" to protect AI companies from being exposed. That's 'boasted' in the past tense because this has since been taken down, in a move to minimize the fact that it's now making a lot of money from its "print books for AI training" business. But not before it was archived. Based on the revelations, it certainly sounds like AI companies are trying to buy up every single book, digitizing them, and then destroying them. This isn't alarmist, but reality. And the reason it's happening, and reportedly accelerating, is that we're living in a post-generative AI world. The value of human-created works, especially in the literary space, is increasing. Untainted by "AI slop," these books that are free from any AI influence are viewed as extremely high-quality training data because they're entirely written by humans.
[4]
Buy, scan, destroy: AI firms are shredding millions of books to train their chatbots
AI companies are buying, cutting apart and scanning millions of physical books to train chatbots, according to a Washington Post report based on newly unsealed court filings. The documents reveal Anthropic's secret "Project Panama" and shed light on similar book-acquisition efforts by Meta, OpenAI and Google, as the AI industry's race for high-quality training data fuels an escalating copyright battle. Artificial intelligence companies are buying millions of physical books, cutting them apart, scanning every page and recycling the remains to train the chatbots powering today's AI boom, according to a report by The Washington Post, based on newly unsealed court filings in a copyright lawsuit against AI startup Anthropic. The filings also reveal how the race to build more capable AI models pushed companies to secure vast collections of books, setting off a wave of copyright lawsuits. The court documents provide one of the clearest glimpses yet into the AI industry's search for high-quality training data. While Anthropic's internal book-scanning project is at the centre of the filings, separate lawsuits involving Meta, OpenAI and Google also highlight how leading AI companies sought access to millions of books to improve their models, although the methods differed across companies. From bookshelf to chatbotAt the centre of the revelations is Anthropic's internal initiative called "Project Panama", which company documents described as an effort to "destructively scan all the books in the world." According to the report, the company spent tens of millions of dollars buying millions of used books, slicing off their spines, scanning every page and sending the remains for recycling to create training data for its Claude AI models. The planning documents also stated: "We don't want it to be known that we are working on this." According to the report, Anthropic initially explored sourcing books from libraries and used bookstores, including New York's Strand Book Store, before eventually purchasing large batches from used-book retailers such as Better World Books and the UK's World of Books. A proposal from one of its scanning vendors said the company planned to digitise between 500,000 and 2 million books in six months, using industrial cutting machines to remove bindings before scanning the pages and sending the remains for recycling. To lead the effort, Anthropic hired Tom Turvey, a former Google executive who helped create Google's Google Books project more than two decades ago, according to The Post. The AI arms raceBooks were viewed as a critical resource because they offered higher-quality writing than much of the internet. One Anthropic co-founder wrote internally that books could teach AI models "how to write well" instead of imitating "low quality internet speak," according to the report. Before launching Project Panama, Anthropic employees had also downloaded books from shadow libraries such as LibGen and Pirate Library Mirror, which host copyrighted material without permission. Anthropic has said it never trained a commercial AI model using the LibGen dataset and did not use Pirate Library Mirror to train any complete AI model. The filings suggest Anthropic was not alone. According to The Post, Meta employees discussed using LibGen to obtain millions of books, with internal messages showing concerns over copyright risks. One engineer reportedly wrote, "Torrenting from a corporate laptop doesn't feel right," while another discussion referred to using rented servers to avoid the activity being traced back to the company. Meta has denied illegally distributing copyrighted works. OpenAI has acknowledged downloading LibGen but told a court it deleted the files before the release of ChatGPT. Google is also facing copyright litigation related to AI training data, with two major publishers recently seeking to join an existing lawsuit against the company. The copyright plot twistThe legal battle over AI training data is far from settled. In June, a US judge ruled that using books to train AI models could qualify as fair use because the process is "transformative." However, the judge also found Anthropic could still face liability over how it acquired some of the books by downloading pirated copies. Anthropic later agreed to pay $1.5 billion to settle claims related to the acquisition of books, while maintaining that the settlement concerned the method of acquisition rather than the legality of AI training itself. Most copyright cases involving AI companies, authors and publishers are still working their way through US courts, meaning the broader legal boundaries for training AI on copyrighted material remain unresolved.
[5]
AI's hunger for data is now consuming rare books and booksellers are cashing in
Artificial intelligence firms are purchasing vast quantities of physical books. These books are scanned to train large language models and then destroyed. Rare and out-of-print titles are particularly sought after for this purpose. Booksellers report a significant surge in demand for these older publications. This practice raises concerns about the loss of scarce literary works. Artificial intelligence companies are increasingly turning to physical books, including rare and antique editions, to feed their large language models, creating a booming market for bulk book purchases that is raising concerns about the destruction of scarce literary works, according to a report by 404 Media. The report says AI companies are buying thousands, and in some cases millions, of used books, scanning their contents into digital datasets and then destroying the physical copies. The practice has gathered momentum as developers seek cleaner, high-quality training material that predates the explosion of AI-generated content online. Also read: Destroyed in 2005, reborn through LNG: Inside America's unlikely energy powerhouse The trend follows legal developments in the United States that have allowed companies to digitise books they legally purchased under the first-sale doctrine. In a lawsuit that has since been settled, court documents showed that Anthropic used hydraulic cutting machines to remove pages from books before scanning them with industrial imaging equipment for AI training. A judge found the digitisation process to be "transformative," making it eligible for fair-use protection because the books were converted into digital form rather than redistributed as new copies. Rare and out-of-print books fuel AI data raceAccording to 404 Media, businesses that once primarily served libraries and booksellers are now actively courting AI firms. One such company, ISBNdb, which describes itself as operating the "world's largest book database," argues that older printed books offer an advantage because they were published before AI-generated text became widespread. "Print books from the pre-LLM era are structurally guaranteed to be free of this contamination. That alone is a significant advantage," the company says on its website, as quoted by 404 Media. It further states: "Physical books published before this date [pre-2022] are structurally clean of modern poisoning tools." According to the report, ISBNdb now facilitates bulk purchases ranging from 1,000 to one million books per order for AI companies. The company also advertises anonymous procurement services, acknowledging the reputational risks surrounding the practice. "The optics problem is real," ISBNdb says on its website. "'AI company destroys two million books' is not a headline that generates sympathy." Also read: In 2006, Canada launched a mission to save Atlantic salmon. Nearly 20 years later, the fish are returning Booksellers see windfall but worry about what's being lostThe surge in demand is already being felt by independent booksellers. One seller told 404 Media that weekly sales jumped from fewer than 20 books to hundreds in April, with buyers selecting seemingly random titles that all carried ISBN numbers, leading the seller to believe AI companies were behind the purchases. The bookseller said many of the works being sold were rare or out of print, raising fears that some of the last remaining physical copies could be disappearing. "I personally have mixed feelings about all of this," the bookseller told 404 Media. "It benefits me financially as well as by clearing out old inventory that is otherwise unlikely to sell. I've been well-suited for these sales with inventory from overseas and foreign language books. On the other hand, I don't like the end-use, and I don't like that uncommon books are being pulped." According to the report, similar buying patterns have been reported by rare-book dealers in the Netherlands, where sellers say they are receiving unusually large orders that they suspect originate from AI companies. While booksellers often cannot verify the identity of buyers because intermediaries such as ISBNdb keep purchasers anonymous, many in the industry believe the bulk acquisitions reflect the AI sector's growing appetite for high-quality, pre-AI training data.
Share
Copy Link
AI companies are purchasing millions of physical books and destroying them through destructive scanning to obtain high-quality human-written training data. The practice targets rare and out-of-print titles, raising ethical concerns about cultural heritage loss while creating a lucrative market for booksellers.

AI companies are aggressively purchasing millions of physical books to train large language models, employing destructive scanning methods that permanently destroy the originals. This practice emerged as developers seek high-quality human-written training data untainted by AI-generated content flooding the internet
1
. The process involves removing book spines, feeding pages through industrial scanners, and recycling or discarding what remains3
.The demand stems from a phenomenon researchers call model collapse—when AI models trained on content produced by earlier models begin reinforcing mistakes while losing variety and less common information from original human data
1
. Physical books printed before the current AI boom offer a dependable record of human writing, providing dense, edited, authoritative content that web crawls cannot replicate.Newly unsealed court filings exposed Anthropic's secretive "Project Panama," an internal initiative described in company documents as an effort to "destructively scan all the books in the world"
4
. The company spent tens of millions of dollars acquiring millions of used books, using hydraulic cutting machines to remove bindings before scanning pages with industrial imaging equipment to train its Claude models4
.Internal planning documents stated: "We don't want it to be known that we are working on this." Anthropic hired Tom Turvey, a former Google executive who helped create Google Books, to lead the effort
4
. One Anthropic co-founder wrote internally that books could teach AI models "how to write well" instead of imitating "low quality internet speak"4
. The company initially explored sourcing from libraries and used bookstores before purchasing large batches from retailers like Better World Books. Proposals indicated plans to digitize between 500,000 and 2 million books in six months4
.ISBNdb, a company operating an online book database, briefly advertised services to source up to 1 million books per order for AI developers before removing the pages following media attention
1
. The company offered books "tailored to your LLM training needs," including older, specialized, rare, and out-of-print titles gathered from used bookstores and catalogs1
. Marketing materials promised a "strict NDA on every engagement," ensuring clients' identities, strategies, and acquisition targets would not be disclosed1
.ISBNdb argued that "print books from the pre-LLM era are structurally guaranteed to be free of this contamination," referring to AI-generated content
5
. The company acknowledged reputational risks, stating on its website: "The optics problem is real. 'AI company destroys two million books' is not a headline that generates sympathy"5
. ISBNdb later told media outlets the service was exploratory and never launched, though the archived pages reveal detailed offerings for bulk book purchases for AI2
.Booksellers across the United States and Europe report receiving unusual bulk orders for obscure and out-of-print titles. Dutch antiquarian bookseller Pieter de Vries received an email from someone identifying as "Natalia" with company "2077AI," requesting quotes for more than 3,000 ISBNs, mostly published between 2020 and 2021 by academic publishers including Emerald Publishing, Elsevier, Wiley, Routledge and Oxford University Press
2
. De Vries initially dismissed it as spam or phishing, and several other Dutch antiquarian booksellers received identical requests2
.One American bookseller told media that weekly sales jumped from fewer than 20 books to hundreds in April, with buyers selecting seemingly random titles that all carried ISBN numbers
5
. The bookseller expressed mixed feelings: "It benefits me financially as well as by clearing out old inventory that is otherwise unlikely to sell. On the other hand, I don't like the end-use, and I don't like that uncommon books are being pulped"5
. Similar buying patterns emerged in Germany and Switzerland, where local media reported secondhand booksellers received bulk orders for highly specialized books inconsistent with traditional collecting2
.The digitization of rare books for AI training has triggered copyright lawsuits against multiple companies. In June, a US federal judge ruled that using books to train AI models could qualify as fair use under the first-sale doctrine because the process is "transformative"—converting books into digital form rather than redistributing them as new copies
4
. However, the judge found Anthropic could still face liability over how it acquired books by downloading pirated copies from shadow libraries like LibGen and Pirate Library Mirror4
.Anthropically later agreed to pay $1.5 billion to settle claims related to book acquisition methods, while maintaining the settlement concerned acquisition rather than the legality of AI training itself
4
. Court filings revealed that before launching Anthropic Project Panama, employees downloaded books from shadow libraries. Anthropic stated it never trained a commercial AI model using the LibGen dataset and did not use Pirate Library Mirror to train any complete AI model4
.Separate lawsuits involving Meta, OpenAI and Google highlight how leading AI companies sought access to millions of books. Meta employees discussed using LibGen with internal messages showing copyright concerns, with one engineer writing: "Torrenting from a corporate laptop doesn't feel right"
4
. OpenAI acknowledged downloading LibGen but told a court it deleted files before releasing ChatGPT4
. Most copyright cases involving AI companies, authors and publishers remain unresolved in US courts, leaving broader legal boundaries for training AI on copyrighted material uncertain4
.Related Stories
Mass book scanning predates current AI applications. Project Gutenberg began creating electronic versions of public-domain works in 1971 and now offers more than 75,000 free ebooks
1
. Google Books launched in 2004, scanning books in partnership with libraries and publishers to create searchable online content. By 2019, Google assembled a collection exceeding 40 million books in over 400 languages1
. Those scans helped form HathiTrust, a digital repository created by research libraries in 20081
.Authors sued Google for scanning copyrighted books without permission, but a federal appeals court ruled in 2015 that the project qualified as fair use
1
. The court determined that creating a searchable index and displaying limited snippets gave books a new purpose without providing readers with a replacement for originals. However, AI companies' bulk book purchases for AI differ from earlier digitization efforts by preservation-focused organizations like the Internet Archive, which has digitized more than 25 million books since 2006 using nondestructive scanning1
.The practice raises ethical concerns about permanently destroying scarce literary works and cultural heritage for AI training data. Rare and out-of-print titles being pulped may represent the last remaining physical copies of certain works
5
. The use of confidentiality agreements and intermediaries suggests AI companies recognize the reputational risks while continuing to pursue high-quality human-written training data through bulk book purchases for AI1
.As the internet increasingly fills with AI-generated content, demand for pre-AI era physical books will likely intensify. Watch for additional legal challenges testing the boundaries of fair use and copyright in AI contexts, potential regulations governing the digitization of rare books, and whether preservation efforts can keep pace with destructive scanning practices. The tension between technological advancement and cultural preservation will shape how AI companies source training data going forward, with implications for both the AI industry and the literary world.
Summarized by
Navi
[3]
23 Jul 2026•Policy and Regulation
26 Jun 2025•Technology

13 Jun 2025•Technology

1
Technology

2
Technology

3
Technology
