Microsoft exec called AI scraping 'largest theft of labor in human history' as court docs expose concerns

Reviewed byNidhi Govil

29 Sources

Share

Newly unsealed court documents in The New York Times copyright lawsuit reveal damaging internal communications from Microsoft and OpenAI executives. Microsoft's Director of Applied Science called AI scraping an astonishing theft while OpenAI's ChatGPT head warned of an existential threat to publishers as click-through rates plummeted by up to 93%.

Microsoft executive labeled AI scraping unprecedented theft

Unredacted court documents from the copyright lawsuit filed by The New York Times against OpenAI and Microsoft have exposed explosive internal communications that undermine the companies' fair use defense

1

. Microsoft Director of Applied Science Brent Hecht described AI scraping in a January 2023 internal memo as "an astonishing theft of unprecedented proportions" and potentially the "largest theft of labor in human history"

2

. In another document, Hecht stated that the plan to widely scrape news content made "a complete mockery of the idea of 'fair use'"

1

. These revelations came from a motion for summary judgment unsealed Thursday, which news plaintiffs led by The New York Times argue demonstrates exactly how OpenAI and Microsoft viewed the threat to journalism before launching AI products like ChatGPT and Copilot

1

.

Source: Futurism

Source: Futurism

OpenAI leadership warned of existential threat to publishers

Internal communications from OpenAI reveal that company executives were acutely aware of the damage their AI training practices would inflict on news organizations. ChatGPT head Nick Turley wrote that publishers would face an "existential threat" from commercial products trained on news content that could substitute news providers

1

. Turley described chatbots as "largely substitutive, period" and predicted they "will get more and more substitutive as they get better"

2

. An OpenAI software engineer acknowledged in internal messages that "no matter how prominently we show the links, users won't click"

1

. OpenAI President Greg Brockman even responded "ah nice" when informed by staffer Nick Ryder about "a hack" to get around The New York Times paywall

3

.

Click-through rates plummeted by up to 93 percent

Microsoft's own data demonstrates the devastating impact of AI products on news organizations. The company recorded 83-93 percent drops in click-through rates for some news plaintiffs, with other publishers experiencing declines between 51-94 percent

1

. Specifically, Microsoft data shows its Copilot "answer engine" caused click-through rates for The New York Times domain to drop as much as 93 percent compared to traditional Bing search

2

. A Microsoft document described this phenomenon as a "doom loop" that would "hurt the performance of our models and the entire web at the same time"

5

. Microsoft CEO Satya Nadella testified under oath that chatbots substitute for news sites by "giving you the information right there on the website on the AI platform versus needing to go to the underlying source"

3

.

Source: TweakTown

Source: TweakTown

Massive scale of content scraping revealed

The unredacted court documents expose the sheer volume of copyrighted material used in AI training practices. OpenAI's mid-training datasets alone contain more than 91,692 copies of works published by The New York Times, Daily News, and Center for Investigative Reporting

2

. A Common Crawl-derived dataset included more than 2 million documents from nytimes.com alone

2

. The filing reveals that "OpenAI delivered the entire GPT-3 training dataset to Microsoft, which Microsoft used to evaluate how to implement OpenAI's models within its own commercial products"

2

. Microsoft similarly provided training data to OpenAI through initiatives called Project Taxi and Project Mango, with Project Mango containing copies of at least 160,903 unique works from news publishers

2

.

Fair use defense faces mounting challenges

The revelations in these unredacted court documents directly contradict the fair use defense that OpenAI and Microsoft have mounted in the copyright infringement lawsuit. Microsoft's own documents acknowledge that "almost no one intended for content they created to be used in this fashion, nor are they compensated for its use"

1

. A Microsoft document stated there is a "real risk" that generative AI could "significantly disrupt the employment of the very people who generated the data on which the foundation model was trained"

2

. News organizations argue they have "compelling evidence of substitution" and believe that proving chatbots are replacing them in their own markets while serving excerpts of articles verbatim will eviscerate the fair use arguments

1

. Microsoft CEO Satya Nadella testified that "anything that is paywalled should be licensed by anyone who wants to use it...for grounding or training" and stated he would have required OpenAI "to retrain its models" if aware they had scraped paywalled content

5

.

Government intervention complicates copyright landscape

The Department of Justice filed a Statement of Interest on September 1, 2026, weighing in on the copyright lawsuit despite not being a party to the case

4

. The DOJ position supports OpenAI and Microsoft, arguing that AI training should be classified as fair use because it is "exceedingly transformative"

4

. The government claims that limiting AI training could hinder US AI development and raise barriers to entry for AI companies to produce large language models

4

. However, this stance conflicts with the experiences of publishers and content creators. Ziff Davis CEO Vivek Shah published commentary arguing that creators deserve licensing payments for their intellectual property

4

. The copyright lawsuit originally filed by The New York Times in late 2023 remains ongoing nearly three years later, with multiple publishers and authors including the Chicago Tribune, The Sun-Sentinel Company, George RR Martin, Michael Chabon, and Sarah Silverman joining related cases

4

.

Source: TechRadar

Source: TechRadar

Economic foundations of journalism under threat

One Microsoft document warned that "it is highly unusual that an end-product threatens the economic foundations of its essential suppliers, but that is the situation we have created for our LLM business with respect to its 'content supply chain'"

2

. News organizations argue that "the future not just of journalism but of responsible AI too depends on preserving incentives for humans to produce the creative works on which a healthy society depends"

1

. The declining news revenue threatens to rob chatbots of the abundant streams of reliable information that makes them valuable tools

1

. Watch for how courts balance transformative use claims against market harm evidence as this case progresses. The outcome will determine whether AI companies must license content or can continue current AI training practices that bypass paywalls and compensation mechanisms. This decision will shape the economic viability of journalism and content creation in an AI-dominated information landscape.

Today's Top Stories

© 2026 TheOutpost.AI All rights reserved