2 Sources
[1]
Here's a balm if the idea of destroying books to train AI breaks your heart
If you can truly appreciate an old book -- and maybe even marvel at how its fragile, yellowing pages contain some of the earliest ways that people tried to make sense of the world around them -- then headlines about tech companies destroying books to train AI likely torture a tender part of your soul. It's indeed depressing to imagine piles of book spines waiting to be fed into wood chippers while torn-out pages are cropped, scanned, and trashed. But that's the cheapest and easiest way to scan books as fast as possible, and AI companies are in a race to advance their models by training on the kind of engaging, high-quality, long-form texts that can only be found in books. So book lovers fear it's likely that the practice is happening on a grander scale than is currently being reported and that some physical copies of books will be lost forever. What makes this destruction extra painful, though, is that it doesn't have to be this way. Google patented a non-destructive book-scanning technology in 2009 that AI firms could use to efficiently scan books -- if they were willing to slow down and invest in the process. It's not perfect, however; studies have found that the curve of the page can distort text, and pages can be missed. As Wired reported, glitches can happen when workers move too quickly, including disembodied hands obscuring pages. Overall, the trade-offs in cost and speed may not appeal to AI firms looking for the cheapest way to scan millions of titles, and Google's method may not be the best way to handle rare books anyway. The Internet Archive, which helps libraries preserve aging collections, has long understood that scanning old texts takes time and attention to limit handling, and that's why it's considered such a human job. For AI companies bent on finding shortcuts, the Internet Archive's method may feel almost alien. But if you're a book lover looking for a balm while parsing rumors of AI-driven book destruction, an old post from 2021 describes how the Internet Archive scans books to avoid damaging even the most precious rare books. In it, a book scanner who has been with the Archive since 2010, Eliza Zhang, offered a moment of zen by explaining what she likes so much about scanning books "the hard way." The Internet Archive did try going the automated route, the post said, even testing out "commercial book scanners that feature a vacuum-powered page-turning arm." But "it turns out those automated scanners didn't really work well for brittle books, rare volumes, and other special collections -- the kinds of material our library partners ask us to digitize," the Internet Archive said. "The job requires keen concentration," Zhang said, since the pages of "very old, fragile books" are "paper thin." In the post, Andrea Mills, who helps lead the Archive's book-scanning operations, explained that "clean, dry human hands are the best way to turn pages." To ensure each rare book only has to go through the scanning process once, Zhang takes her time. She carefully raises the scanner glass with a foot pedal each time she turns a page, then adjusts the cameras and ensures the page is readable. She also takes note of any fold-outs, setting a reminder to go back and scan the inserts so that bonus materials aren't lost while documenting the main pages of the work. The whole time, she understands that if a page is skipped or an image is too blurry, Internet Archive's proprietary software will stop the process and prompt her to scan it again. Practice makes perfect, though, and she reports a low error rate after more than a decade of finding a rhythm in the job. At the time, Zhang had scanned "more than 3 million pages, 14,000 foldouts, and 18,000 items (mostly books)," the post said, with the goal of guaranteeing "zero errors." Ars asked the Internet Archive for comment, but Chris Freeland, the director of library services, said the post detailing Zhang's work is "still the best description of our scanning process today." The post came after a video of Zhang's book scanning got 1.5 million views on what was then Twitter, accompanied by a caption that would resonate with book lovers appalled by AI-driven book destruction today. "At the Internet Archive, this is how we digitize a book," the tweet said. "We never destroy a book by cutting off its binding. Instead, we digitize it the hard way -- one page at a time." Rare booksellers flag suspicious bulk orders Ever since a lawsuit in summer 2025 outed Anthropic for destroying millions of print books to train its AI models, book lovers have moved to defend some of the most precious collections from what feels like AI firms' endless quest to feed all the books in the world into their large language models. The biggest fear for people who want to see books preserved through the training process is that AI firms will callously pulp rare books that can never be replaced. As The Atlantic reported last week, social media "raged" after two recent reports indicated that AI was already endangering rare books. First, a Telegraph report accused Silicon Valley of destroying millions of rare books and "shredding the originals," then 404 Media reported that a book-database company called ISBNdb was advertising that it could help AI firms source books in bulk. This backlash was expected, with ISBNdb reportedly warning its clients that the optics were bad. Anthropic started using a codename -- "Project Panama" -- for its destructive book-scanning in an effort to keep it hidden from the public. There is no indication that Anthropic ever destroyed rare books, and the company has denied doing so in statements. It's also true that some AI firms are helping preserve rare books. OpenAI and Microsoft, for example, are working with Harvard librarians on an initiative to train AI models on about 1 million public-domain books dating back to the 15th century. Others, including Elon Musk's xAI, have publicly said they won't destroy rare books to train AI. In a post on X, Musk said that he "asked the SpaceXAI team to preserve any rare books in a library and scan them the hard way vs just cutting off the spine and scanning." But not everyone is convinced that the biggest AI firms are sincere in their promises to preserve rare books, even if copyright law allows them to destroy copies of books they've purchased. Critics on social media and Reddit have pointed out that Musk didn't say the xAI team would not destroy any books, and it's unclear how his company defines a "rare" book. Meanwhile, rare booksellers continue to flag strange orders, recognizing that most sincere buyers tend to seek out single books, not order large batches of widely varying selections. Just yesterday, an Irish bookstore called Kennys flagged a "bananas" order for 5,000 obscure titles, The Irish Times reported. Unlike bulk orders placed by universities or big libraries, this order seemed suspicious because the buyer made no attempt to haggle on the price -- which has now become an obvious "red flag" that AI is potentially involved, the Times reported. Although some booksellers said they could benefit in the short term from large orders, they fear that selling to AI firms indiscriminately destroying their collections "may be signing their own death warrant," the Times reported. Tomás Kenny of Kennys Bookshop told the Times that AI is "terrifying our industry," and he plans to refuse orders likely destined for AI training in an effort to protect authors' rights to their work and readers' access to hard-to-find titles. "Booksellers tend to be at the forefront of a lot of moral dilemmas and political and ethical issues. In most scenarios, it is not for me to decide," he said. "However, if they are taking a book and scanning it with a view to taking intellectual property, then Kennys do not want to be a part of that." It's possible that public backlash will pressure the biggest firms to adopt methods like the Internet Archive's, which focus as much on preserving rare books physically as they do digitally. But it's heartbreaking to imagine the alternative, where rare books become even rarer, especially for those who love the feeling of holding an old book in their hands. Zhang knows the allure of rare books, occasionally pausing her work to read parts of the titles she scans. When asked what she liked most about her job, Zhang responded with the same gusto that many book lovers feel while browsing the shelves in stores like Kennys. "Everything! I find everything interesting," Zhang said. "I don't feel it is boring. Every collection is important to me."
[2]
Independent bookstores in Europe receive suspicious orders for thousands of books, prompting fears they'll be destroyed to train AI -- sellers believe acquisitions are part of AI tech companies' push to get more data
Some fear that these books will be scanned and destroyed by machines, with no human ever setting their eyes on their content. One independent book retailer in Galway, Ireland, received an online order for 5,000 books, which would be elating for many shop owners, especially at a time when people prefer shopping online from major platforms like Amazon or ditch physical books altogether and buy eBooks instead. However, according to The Irish Times, it's not the number of titles that went into the order that raised red flags, but the obscure books that went with the orders, prompting fears the books are being acquired to train AI, possibly resulting in their destruction. "Some was high-quality non-fiction, like A History of Connemara, and the next thing might be The Eddie Hobbs Guide to your SSIA," Tomás Kenny of Kennys Bookshop told the publication. In this example, the former is a history book that covers a region in western Ireland while the latter is about the Special Savings Incentive Account unique to the country. Berlin bookshops are also reportedly seeing similar orders that contain titles that wouldn't make sense for a person to purchase today, like "Pass Your Driving Test, 2018 Edition." It follows a report last month that AI companies are reportedly shredding millions of books after using them to train AI models, using destructive scanners to quickly digitize the books. Massive orders like these aren't that unique, with many private institutions looking to build libraries often ordering this huge number of books from independent bookstores. However, aside from weird titles included in the order, many large buyers connected with educational and other public institutions often negotiate a bulk price. When this is combined with the weird titles being included in the orders, the sellers cannot help but suspect that these orders are, in fact, made by AI companies looking to ingest more books to add to their training data. The biggest AI companies have been hit with multiple lawsuits regarding book piracy -- for example, court records revealed that Meta torrented 82TB of pirated books for AI training, while Nvidia is in hot water as a judge said that its NeMo Framework "have no other purpose" than to speed up infringement. More recently, Anthropic was hit with a $1.5-billion settlement for infringing the rights of authors and their publishers. Unfortunately, the penalty here is based on the startup's use of pirated books -- the court has ruled that the AI firms' use of these written works to train AI is considered fair use. This isn't applicable in Germany, though. "Under German law scanning books - regardless of the purpose - would not be permissible and constitute a clear violation of copyright law," said Thomas Koch, spokesperson of Germany's Publishers and Booksellers' Association. "That this is now happening with second-hand books is another highly troubling example of this practice." They've also surmised that the purchases are driven by AI bots that troll the internet for ISBNs (the unique numerical code assigned to each book title) that they don't have in their library yet. The orders are shipped to local addresses, but because some European nations have laws that prevent book scanning, the association thinks that these are just collection points for eventual bulk shipping to the U.S. It's not clear if the bookshops are fulfilling these orders or not. On the one hand, these bulk orders are a lifesaver for these stores and could help make surviving society's transition towards eBooks and readers much easier. On the other hand, AI is scaring them, and some fear it wouldn't just be the end of bookstores but could also lead to the degradation of human critical thinking abilities. Follow Tom's Hardware on Google News, or add us as a preferred source, to get our latest news, analysis, & reviews in your feeds.
Share
Copy Link
Independent bookstores across Europe are receiving suspicious bulk orders for thousands of obscure books, raising fears AI companies are acquiring physical books to scan and destroy for AI training data. The practice highlights AI's demand for training data and the ethical dilemma between technological advancement and preserving literary heritage.

Independent bookstores across Europe are receiving alarming bulk orders that have raised concerns about AI companies acquiring books to train AI models. Tomás Kenny of Kennys Bookshop in Galway, Ireland, received an online order for 5,000 books containing unusual combinations of titles, from "A History of Connemara" to "The Eddie Hobbs Guide to your SSIA."
2
Berlin bookshops are seeing similar patterns, with orders including outdated titles like "Pass Your Driving Test, 2018 Edition" that make no sense for typical readers.2
What distinguishes these orders from legitimate institutional purchases is the absence of bulk price negotiations and the seemingly random selection of obscure titles. Booksellers suspect AI bots are trolling the internet for ISBNs—the unique numerical codes assigned to each book title—that aren't yet in their training databases. These orders are shipped to local European addresses, which industry experts believe serve as collection points for eventual bulk shipping to the United States.
2
AI companies are in a race to advance their models by training on high-quality, long-form texts found only in books. The fastest method involves destroying books to train AI through destructive scanners—feeding book spines into wood chippers while torn-out pages are cropped, scanned, and trashed.
1
This practice came to light following a lawsuit in summer 2025 that exposed Anthropic for destroying millions of print books to train its AI models, resulting in a $1.5-billion settlement for copyright infringement.1
2
AI's demand for training data has created an ethical dilemma for book lovers who fear that some physical books, including rare editions, will be lost forever. The concern extends beyond common titles to literary heritage—precious collections that can never be replaced once destroyed. While AI data hunger drives companies to find the cheapest scanning methods, this approach ignores the value of preserving physical books for future generations.
Google patented non-destructive scanning techniques in 2009 that AI companies could use to efficiently scan books without destroying them.
1
However, the method has drawbacks—page curves can distort text, pages can be missed, and glitches occur when workers move too quickly, including disembodied hands obscuring pages. The trade-offs in cost and speed may not appeal to AI firms looking to scan millions of titles as quickly as possible.The Internet Archive demonstrates a more careful approach to preserving books while digitizing them. Eliza Zhang, a book scanner with the Archive since 2010, has scanned more than 3 million pages, 14,000 foldouts, and 18,000 items—mostly books—with a goal of achieving zero errors.
1
The organization tried automated commercial book scanners featuring vacuum-powered page-turning arms but found they didn't work well for brittle books, rare volumes, and special collections. Instead, they discovered that clean, dry human hands are the best way to turn pages and digitize rare books without damage.1
Related Stories
The legal landscape surrounding books to train AI remains complex and varies by jurisdiction. In the United States, courts have ruled that AI firms' use of written works to train AI is considered fair use, though penalties apply for using pirated books. Meta reportedly torrented 82TB of pirated books for AI training, while Nvidia faces legal challenges as a judge stated its NeMo Framework "have no other purpose" than to speed up copyright infringement.
2
German copyright law takes a stricter stance. Thomas Koch, spokesperson of Germany's Publishers and Booksellers' Association, stated: "Under German law scanning books - regardless of the purpose - would not be permissible and constitute a clear violation of copyright law. That this is now happening with second-hand books is another highly troubling example of this practice."
2
This legal protection hasn't stopped suspicious orders from arriving at German bookshops, suggesting AI companies may be circumventing local regulations.Independent bookstores now face an impossible choice. These bulk orders represent financial lifelines that could help them survive society's transition toward eBooks and digital readers. Yet fulfilling these orders may contribute to destroying books to train AI and potentially accelerate the degradation of human critical thinking abilities. The ethical dilemma forces small business owners to weigh immediate economic survival against long-term concerns about AI's impact on literature and culture. Watch for increased regulatory scrutiny in Europe and potential expansion of copyright infringement cases against major AI companies as this practice gains visibility.
Summarized by
Navi
23 Jul 2026•Policy and Regulation
30 Jul 2026•Policy and Regulation

26 Jun 2025•Technology

1
Science and Research

2
Technology

3
Technology
