AI Companies Accused of Destroying Books to Train AI as Independent Bookstores Report Suspicious Orders

2 Sources

Share

Independent bookstores across Europe are receiving suspicious bulk orders for thousands of obscure books, raising fears AI companies are acquiring physical books to scan and destroy for AI training data. The practice highlights AI's demand for training data and the ethical dilemma between technological advancement and preserving literary heritage.

News article

Independent Bookstores Face Suspicious Bulk Orders

Independent bookstores across Europe are receiving alarming bulk orders that have raised concerns about AI companies acquiring books to train AI models. Tomás Kenny of Kennys Bookshop in Galway, Ireland, received an online order for 5,000 books containing unusual combinations of titles, from "A History of Connemara" to "The Eddie Hobbs Guide to your SSIA."

2

Berlin bookshops are seeing similar patterns, with orders including outdated titles like "Pass Your Driving Test, 2018 Edition" that make no sense for typical readers.

2

What distinguishes these orders from legitimate institutional purchases is the absence of bulk price negotiations and the seemingly random selection of obscure titles. Booksellers suspect AI bots are trolling the internet for ISBNs—the unique numerical codes assigned to each book title—that aren't yet in their training databases. These orders are shipped to local European addresses, which industry experts believe serve as collection points for eventual bulk shipping to the United States.

2

AI Tech Companies Push to Get More Data Through Destructive Scanning

AI companies are in a race to advance their models by training on high-quality, long-form texts found only in books. The fastest method involves destroying books to train AI through destructive scanners—feeding book spines into wood chippers while torn-out pages are cropped, scanned, and trashed.

1

This practice came to light following a lawsuit in summer 2025 that exposed Anthropic for destroying millions of print books to train its AI models, resulting in a $1.5-billion settlement for copyright infringement.

1

2

AI's demand for training data has created an ethical dilemma for book lovers who fear that some physical books, including rare editions, will be lost forever. The concern extends beyond common titles to literary heritage—precious collections that can never be replaced once destroyed. While AI data hunger drives companies to find the cheapest scanning methods, this approach ignores the value of preserving physical books for future generations.

Non-Destructive Scanning Techniques Exist But Remain Underutilized

Google patented non-destructive scanning techniques in 2009 that AI companies could use to efficiently scan books without destroying them.

1

However, the method has drawbacks—page curves can distort text, pages can be missed, and glitches occur when workers move too quickly, including disembodied hands obscuring pages. The trade-offs in cost and speed may not appeal to AI firms looking to scan millions of titles as quickly as possible.

The Internet Archive demonstrates a more careful approach to preserving books while digitizing them. Eliza Zhang, a book scanner with the Archive since 2010, has scanned more than 3 million pages, 14,000 foldouts, and 18,000 items—mostly books—with a goal of achieving zero errors.

1

The organization tried automated commercial book scanners featuring vacuum-powered page-turning arms but found they didn't work well for brittle books, rare volumes, and special collections. Instead, they discovered that clean, dry human hands are the best way to turn pages and digitize rare books without damage.

1

Copyright Law and AI Training Data Face Legal Challenges

The legal landscape surrounding books to train AI remains complex and varies by jurisdiction. In the United States, courts have ruled that AI firms' use of written works to train AI is considered fair use, though penalties apply for using pirated books. Meta reportedly torrented 82TB of pirated books for AI training, while Nvidia faces legal challenges as a judge stated its NeMo Framework "have no other purpose" than to speed up copyright infringement.

2

German copyright law takes a stricter stance. Thomas Koch, spokesperson of Germany's Publishers and Booksellers' Association, stated: "Under German law scanning books - regardless of the purpose - would not be permissible and constitute a clear violation of copyright law. That this is now happening with second-hand books is another highly troubling example of this practice."

2

This legal protection hasn't stopped suspicious orders from arriving at German bookshops, suggesting AI companies may be circumventing local regulations.

Bookstores Face Difficult Decisions Amid AI Data Collection

Independent bookstores now face an impossible choice. These bulk orders represent financial lifelines that could help them survive society's transition toward eBooks and digital readers. Yet fulfilling these orders may contribute to destroying books to train AI and potentially accelerate the degradation of human critical thinking abilities. The ethical dilemma forces small business owners to weigh immediate economic survival against long-term concerns about AI's impact on literature and culture. Watch for increased regulatory scrutiny in Europe and potential expansion of copyright infringement cases against major AI companies as this practice gains visibility.

Today's Top Stories

© 2026 TheOutpost.AI All rights reserved