Google faces class action lawsuit from major publishers over alleged AI training copyright violations

Reviewed byNidhi Govil

10 Sources

Share

Major publishers Hachette, Cengage, and Elsevier, along with author Scott Turow, have filed a class action lawsuit against Google in New York federal court. The complaint alleges Google copied millions of copyrighted books and journal articles to train its Gemini AI without permission or compensation, calling it "one of the most prolific infringements of copyrighted materials in history." The suit claims Google violated agreements that allowed limited use of content for services like Google Books.

Major Publishers Sue Google Over Gemini AI Training

A coalition of major publishers and an acclaimed author have launched a class action lawsuit against Google, alleging the tech giant engaged in widespread copyright infringement to train its Gemini AI model. Hachette Book Group, Cengage Learning, Elsevier, and author Scott Turow filed the Google lawsuit on July 10 in the U.S. District Court for the Southern District of New York

1

2

. The complaint describes the alleged infringement as "one of the most prolific infringements of copyrighted materials in history," claiming Google reproduced millions of copyrighted works without permission or compensation

4

.

Source: France 24

Source: France 24

Allegations of Unauthorized Use of Copyrighted Content

The publishers sue Google over claims that the company violated long-standing agreements governing how their content could be used. According to the complaint, publishers and authors historically provided Google with copyrighted works for specific, limited purposes such as making books searchable through Google Books, where users could view only short snippets along with bibliographic information

1

. The plaintiffs allege that Google trained Gemini on copies of these books, as well as books uploaded to Google Play Books and content from Google Scholar, despite never receiving authorization for AI training purposes

4

. "Google illegally copied works from all these scope-limited programs for AI training, knowing it lacked authorization to do so," the lawsuit states

1

.

Source: Gizmodo

Source: Gizmodo

The Google Gemini lawsuit also alleges that the company used web scrapes from "known pirate sources" and accessed content from behind paywalls to amass AI training data

4

. Additionally, the plaintiffs accuse Google of intentionally removing or altering copyright management information on these works to "conceal that its Gemini Models were trained on stolen materials"

1

. This allegation forms the basis of a fourth count under the Digital Millennium Copyright Act, separate from the three counts brought under the Copyright Act

4

.

Internal Documents Suggest Google Knew the Risks

The complaint cites an internal Google document that allegedly warned using copyrighted books for AI training could be "highly problematic for Google" and might result in "$10Bs-$100Bs in potential fines"

1

4

. This evidence suggests Google was aware of the legal risks associated with AI copyright infringement but proceeded anyway. The plaintiffs argue this demonstrates willful infringement, noting that if Google wanted to properly license their content for training purposes, it had the financial resources to do so

5

. The lawsuit points to Alphabet's first $100 billion revenue quarter in October 2025, which the company linked to its AI business, and notes that Gemini has over 650 million monthly active users

4

5

.

How Gemini Competes With Human Authors

The publishers argue that the Gemini AI model now competes directly with the copyrighted works it was trained on, creating outputs that range from near-verbatim copies to replacement textbook chapters and knockoff novels

4

5

. "The scale and speed at which Gemini can create books and compete with human writers is unprecedented, and it can only do that because Google copied plaintiffs' and the class's works to train its AI," the complaint states

2

. The lawsuit provides a striking example: Gemini can produce a 100-page murder mystery in approximately 20 minutes for just $0.39

4

5

. "No publisher or author can compete with that," the filing states

4

.

The suit names specific titles allegedly used in unauthorized AI training, including NK Jemisin's The Fifth Season and Lemony Snicket's Who Could That Be at This Hour?

4

. The publishers also claim that Google has failed to implement effective guardrails to prevent Gemini from producing outputs that substitute for copyrighted works on which it was trained

3

.

Source: Engadget

Source: Engadget

A Growing Wave of AI Training Data Lawsuits

This class action lawsuit against Google joins a mounting number of legal challenges against AI companies over their use of copyrighted material for training. The same group of publishers, including Hachette, Cengage, and Elsevier, along with Scott Turow, filed a similar lawsuit against Meta in May 2026

2

3

. Other AI companies facing copyright challenges include OpenAI, Anthropic, and Meta

1

4

.

While many of these lawsuits remain pending, early court decisions have produced mixed results. Two rulings in California favored AI companies, with judges determining that the use of copyrighted works for AI training constitutes fair use under U.S. copyright law

1

2

. However, both judges emphasized that future cases could reach different conclusions

2

4

. In a notable exception, Anthropic was fined $1.5 billion for pirating works it trained on, marking the largest payout in U.S. copyright law history, with around half a million writers eligible for payments of at least $3,000

1

.

What This Means for the AI Industry

The Southern District of New York venue gives a different judge the opportunity to weigh in on whether AI training constitutes fair use, potentially establishing new precedent outside California

1

4

. The plaintiffs are seeking statutory damages, a permanent injunction, and an order requiring Google to destroy any unauthorized copies used in training

4

. "Copyright law applies to AI companies, including Google, with the same force as every other company that has complied with these laws for decades," the lawsuit asserts

2

5

.

The outcome could reshape how AI companies acquire training data and whether they must compensate rights holders. Some publishers have already opted for licensing deals rather than litigation—HarperCollins signed an agreement with Microsoft in 2024 to provide books for AI training

5

. The plaintiffs allege that Google, which already licenses some content for training, deliberately chose unauthorized sources instead

4

. Google did not respond to requests for comment on the lawsuit

1

4

5

.

Today's Top Stories

© 2026 TheOutpost.AI All rights reserved