OpenAI accused of hiding evidence and lying to court in landmark copyright lawsuit with NYT

Reviewed byNidhi Govil

20 Sources

Share

The New York Times and 16 other publishers are asking a federal judge to sanction OpenAI for allegedly concealing evidence in their copyright infringement case. Court filings claim OpenAI lied for two years about its ability to search training data and ChatGPT logs, while secretly maintaining a database of 78 million conversations that could prove the AI company used copyrighted journalism without permission.

OpenAI Copyright Lawsuit Escalates with Sanctions Motion

The New York Times and a coalition of 16 other publishers filed a sanctions motion Thursday in Manhattan federal court, accusing OpenAI of systematically hiding evidence and lying about its technical capabilities in what could become a landmark copyright infringement case

1

. The allegations center on OpenAI's repeated claims that it lacked the ability to search its systems for copyrighted content, statements now revealed to be false according to testimony from OpenAI privacy engineer Vincent Monaco during an April court-ordered deposition

2

.

The news outlets are requesting legal sanctions including financial penalties and attorney fees, arguing that OpenAI and Microsoft built their AI technologies using millions of news articles without permission

3

. "For over two years, OpenAI lied to The Times, The Daily News Plaintiffs, the public, and the court," said Ian Crosby, lead counsel for the plaintiffs, in a statement

4

.

Source: MediaNama

Source: MediaNama

OpenAI Hid Evidence of 78 Million Chat Logs

Among the most damaging revelations in the ChatGPT copyright trial, Monaco's testimony allegedly disclosed that OpenAI had already amassed a database of approximately 78 million de-identified ChatGPT conversations before The New York Times even filed its lawsuit in 2023

2

. The company was using this data internally to evaluate how much it was infringing on copyrighted works, yet never disclosed its existence to the court or plaintiffs during two years of discovery proceedings

1

.

Additionally, OpenAI allegedly maintained a separate sample of 10 million logs that had already been de-identified and could have been made available early in the discovery process

1

. The publishers claim OpenAI had already searched these samples for New York Times content as part of research into "creating a filter that could be used to block the regurgitation of copyrighted content" through an initiative called "Project Giraffe"

2

.

Source: PYMNTS

Source: PYMNTS

Discovery Misconduct and Unusable Evidence

News outlets urge judge to sanction OpenAI for what they characterize as systematic obstruction during the discovery process. Publishers originally requested a sample of 120 million chat logs, but OpenAI negotiated this down to just 20 million, citing technical burdens and user privacy concerns

2

. When OpenAI finally submitted this smaller sample in December, it had made 19 billion redactions using AI, rendering the evidence "unusable" according to the court

1

.

The plaintiffs further allege that OpenAI deleted billions of relevant ChatGPT conversations after the lawsuit was filed, potentially violating the court's preservation order, and substituted millions of logs in the requested sample

2

. "This motion asks the court to punish OpenAI for hiding and destroying evidence showing how ChatGPT was trained on stolen journalism," said New York Daily News attorney Steven Lieberman

5

.

What's at Stake in the AI Copyright Fight

The OpenAI withholding evidence in copyright lawsuits allegations strike at the heart of how large language models are trained and whether this constitutes fair use under intellectual property law. OpenAI and Microsoft have consistently argued that training AI systems on copyrighted material falls under fair use doctrine, while publishers contend the companies are building "substitutive products" that compete directly with original journalism

5

.

The journalism industry faces mounting pressure from AI-generated search summaries that reduce traffic to original sources. Some data shows small publishers have experienced a 60% traffic drop, with projections suggesting declines exceeding 40% by 2029

3

. The New York Times has already spent more than $28 million fighting AI companies in court, according to financial regulatory filings

5

.

Source: BNN

Source: BNN

The New York Times-led group asks court to sanction OpenAI by preventing the company from using the compromised 20 million log sample as evidence, accepting as fact that chat logs would have shown substantial regurgitation of copyrighted content, and requiring OpenAI to pay legal fees incurred while pursuing improperly withheld evidence

2

.

OpenAI spokesperson Drew Pusateri denied the allegations, stating: "As the Times' case weakens and they've been forced to drop claims against us, they're persisting with their efforts to invade the privacy of people who have nothing to do with this case, including by making these blatantly false allegations"

2

. However, NYT spokesperson Graham James disputed that dropping one claim weakened their case, noting they added claims against Microsoft while maintaining their core allegations .

The outcome could establish critical precedents for how AI companies access and use training data, potentially reshaping the relationship between tech giants and content creators across multiple industries facing similar legal challenges.

Today's Top Stories

© 2026 TheOutpost.AI All rights reserved