AI Companies Face FTC Investigation Over Mass Book Destruction for Model Training

4 Sources

Share

18 advocacy groups have urged the FTC to investigate AI companies like Anthropic and Amazon for buying physical books in bulk, scanning them for AI training data, and then destroying the originals. Critics argue this hoard-and-destroy practice creates unfair barriers to entry and potentially eliminates rare works forever.

News article

Civil Society Groups Push for FTC Action

Eighteen advocacy organizations, including the Demand Progress Education Fund, Consumer Federation of America, and the Institute for Local Self-Reliance, have formally requested the Federal Trade Commission to investigate AI companies for engaging in what they term hoard-and-destroy practices

1

3

. The civil society groups are urging regulators to examine whether AI companies are buying physical books in bulk, digitizing them for AI training data, and then permanently destroying the originals to prevent storage costs and competitive access.

The practice has emerged as a significant concern because it potentially removes nonrenewable resources from broader public access. According to the letter submitted to FTC Chairman Andrew Ferguson and Commissioner Mark Meador, AI companies are engineering a future where only the wealthiest incumbents can build high-quality large language models while operating as sole holders of humanity's written works after destroying the originals

4

. The groups want the FTC to determine whether such conduct constitutes an unfair method of competition under Section 5 of the FTC Act.

Anthropic's Project Panama Revealed

Fresh details about these operations surfaced through court disclosures in Bartz v. Anthropic PBC, a copyright lawsuit brought by authors whose books were used for training without permission

1

. Anthropic's internal initiative, codenamed Project Panama, was described in a 2024 memo as "our effort to destructively scan all the books in the world." The company deliberately kept the operation secret, with internal documents warning employees to avoid discussing it in public areas and not share information with anyone outside Anthropic.

The Washington Post reported in January that Anthropic spent millions to acquire books and physically remove their spines to feed scanned pages into Claude, their AI model

3

. In 2025, a judge ruled that Anthropic's use of legally purchased books to train its AI model did not violate copyright law, accepting the argument that digitizing and destroying a physical book could constitute transformative fair use

4

.

Amazon and Industry-Wide Concerns

Amazon has also been implicated in these practices. A recent investigation by 404 Media revealed that Amazon is buying books in bulk, scanning them to train its AI tools, and then destroying them

4

. Neither Anthropic nor Amazon responded to requests for comment about their book destruction operations.

Intermediaries are feeling pressure from the controversy. ISBNdb, which reportedly helped broker book acquisitions for destructive scanning, recently disavowed the practice and removed a landing page titled "Printed Books Sourcing for Your AI LLMs Dataset Needs," noting they had chosen to pivot away from that direction

1

.

Anticompetitive Behavior and Market Implications

The advocacy groups frame scan-and-destroy operations primarily as anticompetitive behavior that creates barriers to entry for competing AI developers

2

. Kate Oh, special advisor to the Demand Progress Education Fund, stated that "the secretive and reckless way that major AI companies like Anthropic and Amazon are acting shows that there is real smoke here that the FTC needs to investigate."

1

The letter emphasizes that unlike standard data acquisition strategies, this hoard-and-destroy practice could serve as a structural mechanism to raise rival companies' costs and deny startups and fledgling competitors access to source materials essential for competing in the AI marketplace

3

. The groups specifically want the FTC to determine whether dominant AI firms are literally eliminating resources their rivals need to compete, potentially shifting the fight over AI training data from copyright into competition law.

Why Older Books Matter for AI Development

AI companies are particularly targeting older books, especially those published before 2022, because they're unpolluted by AI-generated text that has been seeping into recent written work

1

. These pre-AI era books are guaranteed not to have been written by artificial intelligence and are far more likely to have been well edited, making them valuable training resources for large language models

2

.

The practice has become a public relations headache partly because of the barbarism associated with book burning and its historical connection to authoritarian regimes, and partly because of broader backlash against AI companies for pillaging public resources in pursuit of private gain

1

.

Unanswered Questions About Permanent Loss

A critical concern raised in the letter is whether any texts have been shifted entirely into AI models without leaving any physical copies. The groups acknowledge this cannot be verified from the outside when AI companies refuse to divulge their training data

1

. They're asking the FTC to determine how often destroyed physical books are the last or among the last surviving copies of a given work, particularly concerning rare books that could be lost forever

3

4

.

The FTC under the current administration has carefully balanced being friendly to American business while signaling concern about competition and Big Tech dominance

3

. If regulators proceed with an investigation, it could establish whether these practices constitute monopolization of knowledge and create insurmountable systemic moats around AI incumbents, fundamentally reshaping how AI training data is acquired and regulated.

Today's Top Stories

© 2026 TheOutpost.AI All rights reserved