4 Sources
[1]
Publisher opt-outs of AI training cut Google's DeepMind training data in half.
During its Search antitrust trial yesterday, a DOJ attorney produced a document showing that "80 billion of 160 billion 'tokens' -- snippets of content -- after filtering out the material that publishers had opted out of allowing Google to use for training its AI," according to Bloomberg. But that
[2]
Google's AI Is Scraping Even Sites That Ask to Be Ignored
Don't want a tech conglomerate to train its AI model on your website? Too bad -- Google will do it anyway, thanks to a very convenient workaround. At least, that's more or less what the Silicon Valley behemoth just admitted to in court. As Bloomberg reports, Google said that while it does give
[3]
Google May Train AI on Content for Search Even If Publishers Opt Out
Google Search manages content via the robots.txt web standard Google Search products can reportedly use content from publishers even if they have opted out of artificial intelligence (AI) training. As per the report, a Google DeepMind executive revealed the information during a testimony in the
[4]
Google can train search AI with web content after AI opt-out
During a trial examining Google's search dominance, a Google VP testified that the company trains its AI models on web content, even if publishers opt-out. This data usage, particularly for AI Overviews, raises concerns about revenue loss for publishers. The Justice Department is pushing for
Share
Copy Link
Google's DeepMind VP reveals that the company's search organization can train AI on publisher content even after opt-outs, sparking debates on data usage and monopolistic practices.

In a recent antitrust trial, Google's DeepMind Vice President Eli Collins revealed that the company's search organization can train its AI models on web content even when publishers have opted out of AI training
1
. This admission has sparked concerns about Google's data usage practices and its potential monopolistic behavior in the AI and search markets.Collins confirmed that while publishers can opt out of AI training for DeepMind models, this doesn't extend to other parts of Google, including its search organization
2
. This means that Google's search-specific AI products, such as AI Overviews and the recently launched AI Mode, can still use content from publishers who have opted out of AI training3
.An internal document from 2024 cited during the trial showed that Google had collected 160 billion tokens for AI training data. Half of these tokens were removed due to publisher opt-outs, but based on Collins' testimony, these 80 billion tokens may still be used to train AI within Google's search organization
2
.This revelation has raised concerns about revenue loss for publishers. As Google summarizes answers to search queries using AI at the top of results, users may not click through to independent websites, potentially hurting publishers' ad revenue
4
. The irony is that Google is using data from these same sites to generate AI-powered answers.Google maintains that publishers can manage their content in Search via the robots.txt web standard
1
. However, opting out of being indexed for search is seen as a "death sentence" for websites, effectively leaving publishers with no real choice but to allow their content to be used for AI training2
.Related Stories
The ongoing antitrust case aims to prove that Google has a monopoly in the search and AI space. The U.S. Department of Justice is urging the court to take measures such as forcing Google to sell its Chrome browser, share key search data, and restrict its ability to pay for default search engine status on devices and services
4
.The trial has also revealed Google's exploration of using its vast search data to improve AI models. A document shown in court indicated that Google's CEO of DeepMind, Demis Hassabis, had considered training an AI model with search data, including rankings, to assess the improvement over models not trained with such data
4
.This case highlights the complex relationship between AI development, web content, and publisher rights. It raises questions about the future of AI training practices, the value of web content in the AI era, and the balance between technological advancement and fair competition in the digital landscape.
As the trial continues, the outcome could have significant implications for how tech giants like Google use web data for AI training and potentially reshape the landscape of search and AI technologies.
Summarized by
Navi
22 May 2025•Technology

26 Jun 2026•Business and Economy

10 Aug 2026•Policy and Regulation

1
Technology

2
Technology

3
Policy and Regulation
