20 Sources
[1]
OpenAI faked inability to search training data, hid billions of logs, NYT says
OpenAI is facing calls for "serious sanctions" after fighting to keep news organizations from snooping through millions of logs to find evidence of users skirting their paywalls by prompting ChatGPT to regurgitate their articles. This evidence is considered among the most important to both sides, potentially either dooming OpenAI as an infringer or exonerating its chatbot technology as a transformative fair use of news sites' content. In a sanctions motion Thursday, news organizations suing OpenAI -- led by The New York Times -- accused the AI firm of repeatedly lying for years to conceal evidence of infringement that could hobble OpenAI's defense. These alleged lies were exposed when the court compelled an "ill-prepared witness," OpenAI privacy engineer Vincent Monaco, to be re-deposed. During the subsequent April deposition, he inadvertently revealed that OpenAI misled the court for two years about the cost and burdens of searching ChatGPT logs, NYT's filing said. Among the most shocking revelations, OpenAI allegedly pretended from the earliest stages of the case that it did not have the technical ability to search large anonymized samples of ChatGPT logs when it had actually already conducted such searches prior to the start of litigation, NYT alleged. Sanctions are warranted because "OpenAI's concealment of this fact withheld highly relevant evidence, prolonged discovery, inflated expenses, and burdened the Court," news plaintiffs alleged. Asked for comment, an OpenAI spokesperson suggested that NYT's sanctions motion was a late litigation effort to access more logs and infringe more users' privacy. The spokesperson claimed that when the NYT recently dropped some claims in the lawsuit, it was a sign that news plaintiffs' case was crumbling, not OpenAI's defense. "As the Times' case weakens and they've been forced to drop claims against us, they're persisting with their efforts to invade the privacy of people who have nothing to do with this case, including by making these blatantly false allegations," OpenAI's spokesperson said. "We'll continue defending our users' privacy and the long-established principles of fair use." However, last month, NYT spokesperson Graham James disputed to Ars that news plaintiffs' case was weakened by dropping claims. He suggested instead the suit was streamlined and strengthened by adding claims against Microsoft. "Our core claims remain the same from the day we filed this lawsuit -- that Microsoft and OpenAI stole millions of The Times's copyrighted works to compete with our products and illegally enrich themselves," James said. OpenAI allegedly hid 80M log sample Although the sanctions motion is heavily redacted, it's alleged that Monaco testified that OpenAI had two large samples -- spanning 10 million and 78 million logs -- which had already been de-identified and could have been made available to news plaintiffs early on to maximize the discovery period. "Not once did OpenAI disclose the existence" of those samples over two years, news plaintiffs alleged. Even more frustrating to plaintiffs, OpenAI had already searched those samples for NYT content as part of its research into "creating a filter that could be used to block the regurgitation of copyrighted content," the court filing said. "OpenAI was willing and able to search its output logs -- when it benefitted OpenAI," NYT alleged, accusing the ChatGPT maker of "making the discovery process as burdensome as possible." Court says OpenAI sample is "unusable" In a statement to Ars, NYT's lead counsel, Ian Crosby, suggested that OpenAI obstructed access to logs and distorted evidence to shield its fair use claims. "For over two years, OpenAI lied to The Times, The Daily News Plaintiffs, the public, and the court," Crosby said. "It claimed searching ChatGPT outputs for copies of The Times' and the Daily News Plaintiffs' content was infeasible, burdensome, and invasive of users' privacy -- while at the same time concealing that it had already done such searches. If OpenAI genuinely believed that copying our clients' journalism was fair and legal, it wouldn't have hid the truth about having done it." Instead of being transparent about the existing samples, OpenAI forced news plaintiffs to spend eight months searching in a "sandbox," where they could only access a heavily redacted sample of 20 million logs. That sample was much smaller than the 120 million news logs plaintiffs originally requested, allegedly narrowed due to OpenAI's "false representations regarding its existing technical capabilities" to search larger samples. "This representation is belied by Mr. Monaco's testimony that OpenAI already had the ability" to search "large datasets, such as the more than 80 million output logs," NYT alleged. The 20 million log sample was further "skewed" when OpenAI used AI to make 19 billion redactions to the sample -- so many that the court found the sample "unusable." Eventually, OpenAI removed some of the redactions, but "even then, a large number of redactions remain, including to News Plaintiffs' domains, names, and other fields, which has hampered News Plaintiffs' searches over the data," NYT alleged. Meanwhile, "the entire time that OpenAI was engaging in the improper over-redaction of this sample, it had in its possession a sample of 78 million conversations that had already been de-identified," NYT alleged. "OpenAI did not just oppose production of this evidence based on burden or relevance; it falsely represented to the Court that obtaining this evidence was beyond its capabilities without the expenditure of and months of work and that it would be just as easy for Plaintiffs to do this work -- without disclosing that this work had already been done," news plaintiffs alleged. Similarly frustrating were dragged-out meet-and-confers over data searches that news plaintiffs claimed further limited discovery. For example, very close to discovery ending, OpenAI confusingly claimed that the 78 million log samples had been available for inspection for "over a year," NYT alleged. However, "this makes no sense," news plaintiffs argued, considering OpenAI's very public fight to supposedly defend ChatGPT user privacy by blocking access to any logs beyond the 20 million sample. "Either OpenAI unintentionally produced the dataset and it was so hidden in the training inspection data that even OpenAI did not realize it, or OpenAI knew it buried the dataset in a previous production, but hid that fact from the Court and News Plaintiffs for nearly two years -- all the while vigorously arguing that turning over these logs would violate user privacy," NYT argued. Additionally, news plaintiffs accused OpenAI of other misconduct to obstruct access to evidence. Although the exact amount is redacted, OpenAI randomly deleted some parts of that limited 20 million sample, they alleged. And that's on top of allegedly deleting or compressing billions of logs that should have been preserved. According to NYT, OpenAI's witness testified that OpenAI simply "decided" that complying with the court's sweeping preservation order to retain all chats "would be hard; and thus took no steps to do so." "There can be no question as to the wilfulness of OpenAI's conduct, nor any excuse for its non-compliance. According to Mr. Monaco, OpenAI thought about complying with the Court's Preservation Order, but then decided not to," NYT alleged. "Serious sanctions" necessary News organizations claim that they do not request sanctions against OpenAI "lightly" but that the "severity" of OpenAI's alleged misconduct requires sanctions to punish the AI firm and deter any other AI firms from following a similar playbook. Requesting "severe" sanctions, news plaintiffs want the court to prohibit OpenAI from using the 20 million sample that it fought so hard for. They have further asked the court to find that withheld output logs included "substantial" "regurgitation of News Plaintiffs' copyrighted material" and to block OpenAI from arguing otherwise. Finally, the jury would be instructed that OpenAI deleted billions of logs, which would play into news plaintiffs' narrative that OpenAI has been moving in shady ways to obscure alleged substitution in the market since the case began. "Lesser sanctions would not be effective," news plaintiffs warned. In fact,"serious sanctions are especially appropriate," they said, because OpenAI's misconduct "was knowing and intentional." If the court agrees that OpenAI's misconduct was "egregious," OpenAI's attempt to constrict news organizations' access to logs could end up being a fatal misstep in this intently watched copyright fight. Whether training on copyrighted content is fair use will likely depend on whether news organizations can establish market harms, and OpenAI's defense could be substantially set back if its massively redacted sample is rejected and if that makes it harder to argue substantial infringement did not occur.
[2]
New York Times says OpenAI hid evidence in ChatGPT copyright trial
The New York Times and The Daily News claim that OpenAI has been lying about its ability to search customer chat log data and training datasets for their copyrighted works. It's the latest escalation in a two-year-old lawsuit against the AI firm for allegedly violating copyright law by training its generative AI models on the Times' content and reproducing that journalism in user outputs. Throughout the case, OpenAI has argued that it lacked the ability to search its own training corpus. It also argued that searching or producing its massive collection of ChatGPT conversations would be technically burdensome and would raise user-privacy concerns because the logs would need to be retrieved, processed, and de-identified. The outlets sought that data to determine whether their copyrighted journalism was present in OpenAI's training dataset, and whether and how often ChatGPT generated responses using or reproducing their content. In an April court-ordered deposition, OpenAI data privacy engineer Vinnie Monaco allegedly revealed that OpenAI had already conducted internal searches and evaluations of its training corpus to search for copyrighted journalism works. Monaco's deposition also allegedly revealed that, beginning before the NYT filed its lawsuit, OpenAI had already amassed a database of about 78 million de-identified ChatGPT conversations that it was using internally to determine how much it was infringing on others' works. On top of that dataset, OpenAI also allegedly implemented a so-called "Bloom" filter as part of a set of tools called "Project Giraffe," which detected and kept a record of regurgitation in outputs, shortly after the lawsuit was filed. Those last two revelations are particularly significant. The plaintiffs had originally asked OpenAI to provide a sample of 120 million chat logs, but OpenAI had negotiated to bring the sample down to just 20 million. OpenAI finally submitted that sample to the courts last December, but it had allegedly included so many redactions as to render the sample "unusable," in the court's words. The plaintiffs also claimed OpenAI deleted billions of ChatGPT outputs after they filed suit in direct violation of the court's preservation order, and that the AI giant substituted millions of logs in the requested sample. In other words, they claim OpenAI made it needlessly difficult to obtain information that the company had already collected. "If OpenAI genuinely believed that copying our clients' journalism was fair and legal, it wouldn't have hid the truth about having done it," Ian B. Crosby, lead counsel for the plaintiffs, said in a statement. Now, the NYT and The Daily News are asking the judge to discipline OpenAI for allegedly withholding evidence and messing with the discovery process. They are asking the court to prevent OpenAI from using the 20 million chat log sample as evidence, claiming it is unreliable; to accept as fact that ChatGPT logs would have shown major regurgitation and grounding of the plaintiffs' content; to prevent OpenAI from arguing that its provided chat logs don't demonstrate substantial regurgitation; and to make OpenAI pay legal fees for having to chase down this evidence. In a statement, OpenAI spokesperson Drew Pusateri denied the allegations, accusing the Times of trying to access private user conversations as its case weakens. "As the Times' case weakens and they've been forced to drop claims against us, they're persisting with their efforts to invade the privacy of people who have nothing to do with this case, including by making these blatantly false allegations," Pusateri said. "We'll continue defending our users' privacy and the long-established principles of fair use."
[3]
Publishers Accuse OpenAI of Withholding Evidence in Copyright Lawsuits
On Thursday, multiple news organizations accused OpenAI of withholding evidence about how the company trains its artificial intelligence models in a new motion that's connected to a series of ongoing copyright lawsuits. The motion was filed by 17 publishers, including The New York Times, the New York Daily News, the Chicago Tribune and Ziff Davis (CNET's parent company). Ziff Davis sued OpenAI in 2025, alleging that OpenAI scraped its copyrighted works to train ChatGPT and other large language models. The initial lawsuit dates back to 2023 when The New York Times first sued OpenAI and Microsoft, alleging the companies built their AI technologies using millions of news articles written by journalists. Microsoft and OpenAI have denied the claims. The motion asks the court to impose legal sanctions against OpenAI, but not Microsoft, for allegedly withholding evidence, such as datasets and output logs, and claims that "OpenAI chose obstruction" by failing to produce it. If those sanctions are granted, OpenAI could be ordered to pay financial penalties. "This motion asks the court to punish OpenAI for hiding and destroying evidence showing how ChatGPT was trained on stolen journalism," New York Daily News attorney Steven Lieberman said, per the Associated Press. At the center of the lawsuits is how generative AI, such as ChatGPT, is trained and how it sources its information. The Times' original lawsuit claims that OpenAI's generative AI tools "can generate output that recites Times content verbatim, closely summarizes it, and mimics its expressive style," raising questions of copyright infringement. The lawsuits come amid a broader conversation in the journalism industry: declining traffic across digital media outlets. AI overviews are often cited as a major reason for the decline in clicks to original reporting by writers and editors, which in turn impacts publishers' advertising revenue. A growing reliance on AI chatbots for finding news and other content is also a major concern for publishers, as it siphons off loyal readership and audience. Some data shows that small publishers have been hit the hardest, with a reported 60% traffic drop, while another analysis predicts traffic declines of more than 40% by 2029. A statement by Ziff Davis notes that "OpenAI has copied and monetized Ziff Davis content without permission on a massive scale." Lance Koonce, partner at Klaris Law and counsel for Ziff Davis, said that, since the lawsuit, "OpenAI repeatedly lied about its ability to search its own data sets for Ziff Davis content and engaged in other serious litigation misconduct." An ongoing debate over copyright and AI OpenAI has long maintained that AI training is fair use. An OpenAI spokesperson denied the allegations in a statement to CNET, stating: "As the Times' case weakens and they've been forced to drop claims against us, they're persisting with their efforts to invade the privacy of people who have nothing to do with this case, including by making these blatantly false allegations." The statement went on to say: "We'll continue defending our users' privacy and the long-established principles of fair use." In a 2024 rebuttal to the original lawsuit filed by The New York Times, OpenAI said the publisher falsely accused the company of destroying data and instead accused the newspaper of "secretly" deleting its own data that would have shown internal use of OpenAI products. Although the Times has dropped one claim against OpenAI, the larger lawsuit remains in litigation. Other tech giants, including Meta, have also been accused by authors and news publishers of copyright infringement. Many of those cases are still in litigation as courts decide where to draw the line between fair use and infringement in the age of AI.
[4]
New York Times-led group asks court to sanction OpenAI in US copyright dispute
July 9 (Reuters) - A group of newspapers including the New York Times (NYT.N), opens new tab and New York Daily News asked a federal court in Manhattan on Thursday to sanction OpenAI in their high-stakes copyright dispute for allegedly lying to the court about its ability to search its systems for proof that it misused millions of their articles in AI training. The newspapers told the court, opens new tab in a filing that OpenAI falsely told the court it could not search its large language models for their copyrighted material while hiding that it had done so "even before the first News Plaintiff filed suit." The newspapers said that OpenAI had also deleted billions of relevant ChatGPT conversations or made them unsearchable. They asked the court for sanctions, including attorneys' fees, and a court finding that OpenAI's chat logs showed that the company misused their copyrighted works. Spokespeople for OpenAI did not immediately respond to a request for comment on the motion. The lawsuit, first filed by the Times in 2023, accused OpenAI and its largest financial backer Microsoft (MSFT.O), opens new tab of using millions of its articles without permission to train the large language model behind OpenAI's popular chatbot ChatGPT. The case is one of many brought by copyright owners including authors, visual artists and music labels against tech companies such as OpenAI, Anthropic and Meta Platforms for allegedly misusing their material to train AI systems. "For over two years, OpenAI lied to The Times, The Daily News Plaintiffs, the public, and the court," the New York Times' lead attorney Ian Crosby said in a statement. "It claimed searching ChatGPT outputs for copies of The Times' and the Daily News Plaintiffs' content was infeasible, burdensome, and invasive of users' privacy - while at the same time concealing that it had already done such searches." OpenAI previously told the court that it did not have tools to search its datasets and output logs for copyrighted material, but an OpenAI employee later testified that the company had "performed multiple searches for News Plaintiffs' content," according to the newspapers' Thursday filing. Reporting by Blake Brittain in Washington; edited by David Gaffen Our Standards: The Thomson Reuters Trust Principles., opens new tab * Suggested Topics: * Litigation * Data Privacy * Intellectual Property Blake Brittain Thomson Reuters Blake Brittain reports on intellectual property law, including patents, trademarks, copyrights and trade secrets, for Reuters Legal. He has previously written for Bloomberg Law and Thomson Reuters Practical Law and practiced as an attorney.
[5]
News outlets urge a judge to sanction OpenAI in a high-stakes AI copyright fight
NEW YORK (AP) -- The New York Times, the Daily News and other media outlets are asking a federal judge to impose sanctions on OpenAI, escalating a fight over artificial intelligence and copyright that could shape the future of a struggling news industry. The newspapers allege the ChatGPT maker is hiding evidence important to what could be a landmark copyright infringement trial over how OpenAI and its business partner, Microsoft, built their AI technologies using millions of news articles. At issue is whether AI chatbots are unfairly competing as an information source, siphoning off web traffic without doing the journalistic work involved in gathering the news. A filing Thursday in a Manhattan federal courthouse alleges OpenAI "chose obstruction" over releasing datasets and ChatGPT logs that could show how the AI system used copyrighted news content. The plaintiffs are asking the judge to penalize the company for "discovery misconduct" that could distort evidence, saying a recent deposition of an OpenAI employee contradicts the company's earlier claims. New York Daily News attorney Steven Lieberman said OpenAI has been "making misrepresentations" for two years about its ability to search for copyrighted content in its AI training datasets and logs. "This motion asks the court to punish OpenAI for hiding and destroying evidence showing how ChatGPT was trained on stolen journalism," said Lieberman, who represents the Daily News and seven of its sister papers. OpenAI didn't immediately respond to a request for comment Thursday. The New York Times sued OpenAI and Microsoft in late 2023, about a year after ChatGPT's debut sparked a commercial AI boom and began changing the way people search for information online. The threat to news publications became even more apparent when Google in 2024 introduced AI-generated summaries at the top of online search results, cutting off the advertising dollars that come when people click a link to the information's original source. The Times has since been joined by other news organizations, including Daily News and Chicago Tribune parent MediaNews Group, digital media publisher Ziff Davis and the nonprofit Center for Investigative Reporting. OpenAI and other tech companies have argued the process of training their AI systems on digitized books, online articles and other writings found on the internet is protected by the "fair use" doctrine of U.S. copyright law. It's a theory being tested in dozens of lawsuits as visual artists, novelists, music record labels and other creative industries take AI companies to court, with mixed results. In the case involving the biggest copyright settlement so far, OpenAI rival Anthropic agreed to pay book authors $1.5 billion for training its chatbot Claude on their pirated works -- an amount that represents a small fraction of Anthropic's $965 billion market valuation as it prepares to become publicly traded. The New York Times' arguments are different from those brought by book authors. In its original lawsuit and an amended complaint filed last month, it focused on the unfair competition of companies that "seek to free-ride on The Times's massive investment in its journalism by using it to build substitutive products without permission or payment." The Times has already spent more than $28 million on fighting AI companies in court, according to filings with financial regulators that disclose its litigation costs. The costs include another lawsuit the newspaper filed last year against AI company Perplexity. Among the sanctions sought by the newspapers Thursday are attorney fees that would pay for the efforts to secure "improperly withheld" evidence. The mounting costs come as a growing number of media organizations have signed licensing deals with OpenAI and other AI companies such as Google and Facebook parent Meta that typically pay the outlet a fee to be able to train AI systems on their news feeds or archives. The Associated Press was the first to announce such a deal with OpenAI in 2023. ___ O'Brien reported from Providence, Rhode Island.
[6]
New York Times and Other Publishers Ask Court to Penalize OpenAI
The Times, The New York Daily News and other media organizations accused OpenAI of withholding evidence in a lawsuit. The New York Times, The New York Daily News and 15 other media organizations said in a federal court filing on Thursday that OpenAI was withholding evidence that could play a key role in high-profile lawsuits the companies filed against the artificial intelligence start-up. With their filing, the publishers called for legal sanctions against OpenAI, accusing the company of violating court rules and acting in bad faith during the litigation's fact discovery phase. The Times sued OpenAI in late 2023, accusing the company of infringing on its copyrights by using its materials to train ChatGPT and other technologies. In the months that followed, other publishers sued the A.I. start-up, making similar accusations. Many of those cases were consolidated last year. In the motion on Thursday, known as sanctions filing, the publishers said OpenAI refused to provide information showing how the company's A.I. systems are trained and used. "The evidence is in OpenAI's training data sets and ChatGPT output logs," the parties said in their motion. "But instead of just producing that evidence at the start of the case and focusing on the merits of its fair use defense, OpenAI chose obstruction." OpenAI did not immediately respond to a request for comment. OpenAI has previously denied wrongdoing, saying it respects the rights of content creators. The company has also argued in a court filing that ChatGPT is not a substitute for a Times subscription. A sanctions filing is an unusual step that forces a judge to settle a legal disagreement, said Robin Feldman, a professor at U.C. Law San Francisco. It "makes the judge get down in the mud with other parties," she added. The Times was the first major American media company to sue OpenAI over copyright issues related to its written works. The Times's suit made similar accusations against Microsoft, one of OpenAI's primary partners. Microsoft has denied the allegations. A judicial panel last year consolidated many of the dozens of cases brought by publishers against OpenAI, including the lawsuits from The Times and from authors, including the comedian Sarah Silverman, John Grisham, Jonathan Franzen and George R.R. Martin. Like other A.I. companies, OpenAI has built its technologies by feeding them enormous amounts of data, some of which is copyrighted. OpenAI, Microsoft and other companies have long claimed that they can legally use copyrighted material to train their A.I. systems without paying for it because they transformed the material for a different use. The Times, The Daily News and other publishers filed their motion after deposing an OpenAI employee. The deposition, which was largely redacted in public court documents, shows that OpenAI could have provided the data the plaintiffs have long sought, the publishers claimed. "For two years, OpenAI has been making misrepresentations to the court regarding its ability to search for Daily News content in its training data sets and output logs," said Steven Lieberman, counsel for The Daily News and several other newspapers that have sued OpenAI. "OpenAI lied to The Times, the Daily News plaintiffs, the public and the court," said Ian B. Crosby, a partner at Susman Godfrey and the lead counsel for The Times. The publishers are asking for monetary penalties and other sanctions, according to the filing. The filing does not ask for sanctions against Microsoft.
[7]
News outlets ask judge to sanction OpenAI in copyright case
The New York Times, the Daily News, and others accuse OpenAI of hiding evidence in a case that could reshape how AI is built, and how journalism survives A group of news publishers has asked a federal judge to impose sanctions on OpenAI. The New York Times, the Daily News, and others allege the ChatGPT maker is concealing evidence central to their copyright case, the Associated Press reports. A filing on Thursday in Manhattan federal court claims OpenAI "chose obstruction" over handing over datasets and ChatGPT logs. Those records could show how the system used copyrighted news content to train. The publishers accuse OpenAI of "discovery misconduct", saying a recent deposition of an OpenAI employee contradicts the company's earlier claims. Daily News lawyer Steven Lieberman said OpenAI had spent two years "making misrepresentations" about its ability to search its training data. The motion asks the court to punish OpenAI for hiding and destroying evidence, in Lieberman's words. OpenAI did not immediately respond to a request for comment. The stakes reach well beyond one filing. The Times sued OpenAI and Microsoft in late 2023, and has since been joined by a wave of other newspapers, alongside Ziff Davis and the Center for Investigative Reporting. Fair use, or free-riding At the heart of the fight is a simple question with no settled answer. OpenAI argues that training AI on public writing is protected by copyright's "fair use" doctrine, a defence being tested in dozens of suits from artists, novelists, and music labels. The Times frames it differently, as unfair competition. It says AI firms free-ride on its costly journalism to build "substitutive" products that answer readers without sending them, or ad money, back to the source. That threat sharpened when AI-generated search answers began cutting publisher traffic. Courts are only starting to weigh in, with a German court finding Google liable for its AI Overviews. A costly, forking road The litigation is expensive. The Times says it has spent more than $28m fighting AI companies, including a separate suit against Perplexity, and now wants OpenAI to cover fees for chasing withheld evidence. There is a benchmark for what losing can cost. Anthropic agreed to pay book authors $1.5bn, roughly $3,000 per work, a landmark sum that still amounts to a sliver of its valuation. Not everyone is suing, though. Many outlets have signed licensing deals with AI firms, and even Getty Images struck a pact with a company it had sued, while regulators pursue their own remedies, such as France's €250m fine against Google. That split, sue or license, is the industry's central bet on its own future. A sanctions ruling against OpenAI would not settle the copyright question, but it could hand publishers leverage they have so far struggled to find.
[8]
New York Times, publishers seek sanctions against OpenAI
The New York Times, The New York Daily News, and 15 other media organizations filed a motion Thursday in a Manhattan federal court asking a judge to impose sanctions on OpenAI, alleging the company withheld and destroyed evidence central to a high-stakes copyright lawsuit. At the heart of the publishers' complaint is an allegation that OpenAI misled the court by asserting it lacked the capability to search its systems for copyrighted material -- a claim the publishers say was false, given that OpenAI had conducted exactly such searches before any of the news organizations initiated legal action. Separately, the publishers contend that OpenAI rendered billions of ChatGPT exchanges inaccessible, either through deletion or by stripping them of searchability. "Instead of just producing that evidence at the start of the case and focusing on the merits of its fair use defense, OpenAI chose obstruction," the motion stated. The filing came after the plaintiffs deposed an OpenAI employee whose testimony, they say, contradicts the company's earlier representations to the court. According to Steven Lieberman, who serves as counsel for The Daily News and a number of co-plaintiff papers, the company spent two years misleading the court about whether it could locate copyrighted material within its training datasets and output logs -- a characterization he framed using the word "misrepresentations." Ian B. Crosby, lead counsel for The Times, said OpenAI "lied to The Times, the Daily News plaintiffs, the public and the court." OpenAI spokesperson Drew Pusateri denied the allegations. "As the Times' case weakens and they've been forced to drop claims against us, they're persisting with their efforts to invade the privacy of people who have nothing to do with this case, including by making these blatantly false allegations," Pusateri said in a statement. "We'll continue defending our users' privacy and the long-established principles of fair use." User privacy has been the rationale OpenAI has cited in prior proceedings to justify its reluctance to turn over ChatGPT conversation logs. The publishers are seeking monetary penalties, attorney fees, and a court finding that OpenAI's chat logs showed the company misused their copyrighted works. The litigation traces back to a complaint the Times lodged against OpenAI and Microsoft $MSFT in late 2023, alleging that the two companies built ChatGPT in part by ingesting millions of Times articles without seeking permission or offering compensation. The case has since been joined by other news organizations, including MediaNews Group-owned papers and digital publisher Ziff Davis, with many suits consolidated by a judicial panel. The Times has spent more than $28 million on AI-related litigation, according to filings with financial regulators. Across the broader AI copyright landscape, OpenAI and its peers have staked their defense on the principle that ingesting written material to train machine-learning models constitutes fair use under federal copyright law -- an argument that courts are weighing in a wave of suits filed by novelists, visual artists, and the music industry.
[9]
News outlets seek sanctions against OpenAI in copyright battle
OpenAI has been "hiding and destroying evidence" of how it trained ChatGPT on copyrighted news content, US media organisations allege as legal costs in the landmark copyright battle top $28 million. Media organisations including the New York Times and the Daily News are asking a federal judge to impose sanctions on OpenAI, escalating a legal fight over artificial intelligence and copyright that could reshape the future of a struggling news industry. The newspapers allege the ChatGPT maker is concealing evidence central to what could be a landmark copyright infringement trial over how OpenAI and its business partner, Microsoft, built their AI systems using millions of news articles. At stake is whether AI chatbots are unfairly competing as an information source, draining web traffic without doing the journalistic work involved in gathering the news. A filing on Thursday in a Manhattan federal court alleges OpenAI "chose obstruction" over releasing datasets and ChatGPT logs that could show how the AI system used copyrighted news content. The plaintiffs are asking the judge to penalise the company for "discovery misconduct" that could distort evidence, saying a recent deposition of an OpenAI employee contradicts the company's earlier claims. New York Daily News attorney Steven Lieberman said OpenAI had been "making misrepresentations" for two years about its ability to search for copyrighted content in its AI training datasets and logs. "This motion asks the court to punish OpenAI for hiding and destroying evidence showing how ChatGPT was trained on stolen journalism," said Lieberman, who represents the Daily News and seven of its sister papers. The New York Times sued OpenAI and Microsoft in late 2023, about a year after ChatGPT's debut sparked a commercial AI boom and began changing the way people search for information online. The threat to news publications became more acute in 2024, when Google introduced AI-generated summaries at the top of search results, cutting off the advertising revenue generated when readers click through to an original source. The Times has since been joined by other news organisations, including Daily News and Chicago Tribune parent MediaNews Group, digital media publisher Ziff Davis and the nonprofit Center for Investigative Reporting. OpenAI and other tech companies have argued that training their AI systems on digitised books, online articles and other web content is protected by the "fair use" doctrine of US copyright law -- a theory being tested in dozens of lawsuits as visual artists, novelists, music labels and other creative industries take AI companies to court, with mixed results. In the largest copyright settlement so far, OpenAI rival Anthropic agreed to pay book authors $1.5 billion (€1.35bn) for training its Claude chatbot on their works without authorisation. The Times's arguments differ from those brought by book authors. In its original lawsuit and an amended complaint filed last month, it focused on the unfair competition of companies that seek to profit from its journalism without permission or payment to build rival products. The Times has already spent more than $28 million (€25m) fighting AI companies in court, according to regulatory filings disclosing its litigation costs -- including a separate lawsuit filed last year against AI company Perplexity. Among the sanctions sought on Thursday are attorney fees to cover the cost of securing what the newspapers call "improperly withheld" evidence. The escalating legal costs come as a growing number of media organisations have signed licensing deals with OpenAI and other AI companies, including Google and Meta, that pay outlets a fee to train AI systems on their news feeds or archives.
[10]
The New York Times is escalating its fight with OpenAI, urging a judge to impose sanctions
News outlets are asking a federal judge to sanction OpenAI, claiming it's hiding ChatGPT logs relevant to a landmark copyright trial. The New York Times, the Daily News and other media outlets are asking a federal judge to impose sanctions on OpenAI, escalating a fight over artificial intelligence and copyright that could shape the future of a struggling news industry. The newspapers allege the ChatGPT maker is hiding evidence important to what could be a landmark copyright infringement trial over how OpenAI and its business partner, Microsoft, built their AI technologies using millions of news articles. At issue is whether AI chatbots are unfairly competing as an information source, siphoning off web traffic without doing the journalistic work involved in gathering the news. A filing Thursday in a Manhattan federal courthouse alleges OpenAI "chose obstruction" over releasing datasets and ChatGPT logs that could show how the AI system used copyrighted news content. The plaintiffs are asking the judge to penalize the company for "discovery misconduct" that could distort evidence, saying the recent deposition of an OpenAI employee contradicts the company's earlier claims. New York Daily News attorney Steven Lieberman said OpenAI has been "making misrepresentations" for two years about its ability to search for copyrighted content in its AI training datasets and logs. "This motion asks the court to punish OpenAI for hiding and destroying evidence showing how ChatGPT was trained on stolen journalism," said Lieberman, who represents the Daily News and seven of its sister papers. OpenAI has described its limitations in sharing ChatGPT logs as a measure to protect user privacy. "As the Times' case weakens and they've been forced to drop claims against us, they're persisting with their efforts to invade the privacy of people who have nothing to do with this case, including by making these blatantly false allegations," said a statement Thursday from OpenAI spokesperson Drew Pusateri. "We'll continue defending our users' privacy and the long-established principles of fair use." The New York Times sued OpenAI and Microsoft in late 2023, about a year after ChatGPT's debut sparked a commercial AI boom and began changing the way people search for information online. The threat to news publications became even more apparent when Google in 2024 introduced AI-generated summaries at the top of online search results, cutting off the advertising dollars that come when people click a link to the information's original source. The Times has since been joined by other news organizations, including MediaNews Group-owned newspapers the Daily News and the Chicago Tribune, digital media publisher Ziff Davis and the nonprofit Center for Investigative Reporting. OpenAI and other tech companies have argued the process of training their AI systems on digitized books, online articles and other writings found on the internet is protected by the "fair use" doctrine of U.S. copyright law. It's a theory being tested in dozens of lawsuits as visual artists, novelists, music record labels and other creative industries take AI companies to court, with mixed results. In the case involving the biggest copyright settlement so far, OpenAI rival Anthropic agreed to pay book authors $1.5 billion for training its chatbot Claude on their pirated works -- an amount that represents a small fraction of Anthropic's $965 billion market valuation as it prepares to become publicly traded. The New York Times' arguments are different from those brought by book authors. In its original lawsuit and an amended complaint filed last month, it focused on the unfair competition of companies that "seek to free-ride on The Times's massive investment in its journalism by using it to build substitutive products without permission or payment." The Times has already spent more than $28 million on fighting AI companies in court, according to filings with financial regulators that disclose its litigation costs. The costs include another lawsuit the newspaper filed last year against AI company Perplexity. Among the sanctions sought by the newspapers Thursday are attorney fees that would pay for the efforts to secure "improperly withheld" evidence. The mounting costs come as a growing number of media organizations have signed licensing deals with OpenAI and other AI companies such as Google and Facebook parent Meta that typically pay the outlet a fee to be able to train AI systems on their news feeds or archives. The Associated Press was the first to announce such a deal with OpenAI in 2023.
[11]
News outlets ask court to sanction OpenAI in copyright case
Ranking Democrats' most winnable Senate races after Platner's withdrawal | DATA NERDS Several news outlets, including The New York Times, asked a judge Thursday to sanction OpenAI for allegedly withholding evidence in a copyright dispute with the ChatGPT maker. The newspapers, which have accused the AI firm of improperly using their copyrighted work to train its models, argue that OpenAI incorrectly claimed it could not search and preserve its training data and output logs. "This is a case about copying," they wrote in the filing. "There is no question that it happened. Nor should there be one about what was copied, how often, or to what end. The evidence is in OpenAI's training datasets and ChatGPT output logs." "But instead of just producing that evidence of the start of the case and focusing on the merits of its fair use defense, OpenAI chose obstruction," they continued. The New York Times sued OpenAI and Microsoft, which was a key early investor in the company, for copyright infringement in December 2023. Other outlets brought similar lawsuits, which have since been consolidated into a single case. They allege OpenAI "intentionally hid its discovery capabilities from News Plaintiffs and the Court for two years." The newspapers are asking the court to find that the outputs from the company's chatbot show or would have shown "substantial and systematic grounding on and regurgitation" of their copyrighted work, in addition to seeking attorneys' fees. The Hill has reached out to OpenAI for comment. The company has previously argued that training its models on publicly available content is covered by fair use, a legal principle that allows for the use of copyrighted materials without permission in particular cases. The ChatGPT maker has also suggested that pure regurgitation of newspapers' work by the chatbot is a rare occurrence. While the copyright dispute plays out in court, some outlets have instead reached licensing agreements with major AI companies. The Associated Press struck an agreement with OpenAI in 2023, as did Politico's parent company, Axel Springer. News Corp, which owns The Wall Street Journal, similarly announced a partnership with the AI firm in 2024.
[12]
News Outlets Urge a Judge to Sanction OpenAI in a High-Stakes AI Copyright Fight
NEW YORK (AP) -- The New York Times, the Daily News and other media outlets are asking a federal judge to impose sanctions on OpenAI, escalating a fight over artificial intelligence and copyright that could shape the future of a struggling news industry. The newspapers allege the ChatGPT maker is hiding evidence important to what could be a landmark copyright infringement trial over how OpenAI and its business partner, Microsoft, built their AI technologies using millions of news articles. At issue is whether AI chatbots are unfairly competing as an information source, siphoning off web traffic without doing the journalistic work involved in gathering the news. A filing Thursday in a Manhattan federal courthouse alleges OpenAI "chose obstruction" over releasing datasets and ChatGPT logs that could show how the AI system used copyrighted news content. The plaintiffs are asking the judge to penalize the company for "discovery misconduct" that could distort evidence, saying a recent deposition of an OpenAI employee contradicts the company's earlier claims. New York Daily News attorney Steven Lieberman said OpenAI has been "making misrepresentations" for two years about its ability to search for copyrighted content in its AI training datasets and logs. "This motion asks the court to punish OpenAI for hiding and destroying evidence showing how ChatGPT was trained on stolen journalism," said Lieberman, who represents the Daily News and seven of its sister papers. OpenAI didn't immediately respond to a request for comment Thursday. The New York Times sued OpenAI and Microsoft in late 2023, about a year after ChatGPT's debut sparked a commercial AI boom and began changing the way people search for information online. The threat to news publications became even more apparent when Google in 2024 introduced AI-generated summaries at the top of online search results, cutting off the advertising dollars that come when people click a link to the information's original source. The Times has since been joined by other news organizations, including Daily News and Chicago Tribune parent MediaNews Group, digital media publisher Ziff Davis and the nonprofit Center for Investigative Reporting. OpenAI and other tech companies have argued the process of training their AI systems on digitized books, online articles and other writings found on the internet is protected by the "fair use" doctrine of U.S. copyright law. It's a theory being tested in dozens of lawsuits as visual artists, novelists, music record labels and other creative industries take AI companies to court, with mixed results. In the case involving the biggest copyright settlement so far, OpenAI rival Anthropic agreed to pay book authors $1.5 billion for training its chatbot Claude on their pirated works -- an amount that represents a small fraction of Anthropic's $965 billion market valuation as it prepares to become publicly traded. The New York Times' arguments are different from those brought by book authors. In its original lawsuit and an amended complaint filed last month, it focused on the unfair competition of companies that "seek to free-ride on The Times's massive investment in its journalism by using it to build substitutive products without permission or payment." The Times has already spent more than $28 million on fighting AI companies in court, according to filings with financial regulators that disclose its litigation costs. The costs include another lawsuit the newspaper filed last year against AI company Perplexity. Among the sanctions sought by the newspapers Thursday are attorney fees that would pay for the efforts to secure "improperly withheld" evidence. The mounting costs come as a growing number of media organizations have signed licensing deals with OpenAI and other AI companies such as Google and Facebook parent Meta that typically pay the outlet a fee to be able to train AI systems on their news feeds or archives. The Associated Press was the first to announce such a deal with OpenAI in 2023. ___ O'Brien reported from Providence, Rhode Island.
[13]
New York Times-led group asks court to sanction OpenAI in US copyright dispute
Newspapers have asked a federal court to sanction OpenAI for allegedly lying about its AI training data. They claim OpenAI falsely stated it could not search its systems for copyrighted articles. This comes as OpenAI is accused of misusing millions of news articles without permission. The newspapers seek attorneys' fees and a court finding of misuse of their works. A group of newspapers including the New York Times and New York Daily News asked a federal court in Manhattan on Thursday to sanction OpenAI in their high-stakes copyright dispute for allegedly lying to the court about its ability to search its systems for proof that it misused millions of their articles in AI training. The newspapers told the court in a filing that OpenAI falsely told the court it could not search its large language models for their copyrighted material while hiding that it had done so "even before the first News Plaintiff filed suit." The newspapers said that OpenAI had also deleted billions of relevant ChatGPT conversations or made them unsearchable. They asked the court for sanctions, including attorneys' fees, and a court finding that OpenAI's chat logs showed that the company misused their copyrighted works. Spokespeople for OpenAI did not immediately respond to a request for comment on the motion. The lawsuit, first filed by the Times in 2023, accused OpenAI and its largest financial backer Microsoft of using millions of its articles without permission to train the large language model behind OpenAI's popular chatbot ChatGPT. The case is one of many brought by copyright owners including authors, visual artists and music labels against tech companies such as OpenAI, Anthropic and Meta Platforms for allegedly misusing their material to train AI systems. "For over two years, OpenAI lied to The Times, The Daily News Plaintiffs, the public, and the court," the New York Times' lead attorney Ian Crosby said in a statement. "It claimed searching ChatGPT outputs for copies of The Times' and the Daily News Plaintiffs' content was infeasible, burdensome, and invasive of users' privacy - while at the same time concealing that it had already done such searches." OpenAI previously told the court that it did not have tools to search its datasets and output logs for copyrighted material, but an OpenAI employee later testified that the company had "performed multiple searches for News Plaintiffs' content," according to the newspapers' Thursday filing.
[14]
NYT-Led Publishers Escalate Copyright Fight, Seek Court Sanctions Against OpenAI
The New York Times and several other news organizations have intensified their copyright lawsuit against OpenAI, asking a federal judge in Manhattan to sanction the company over allegations that it misled the court about its ability to identify copyrighted material within its AI systems. According to court filings cited by AP News, the publishers argue OpenAI withheld datasets and ChatGPT usage records that are critical to determining whether their copyrighted articles were used in ways that violate copyright law. "OpenAI misled the Court and News Plaintiffs regarding its existing technical capacity all while continuing to compress and delete billions of conversation logs," the plaintiffs said in the filing. "As the Times' case weakens and they've been forced to drop claims against us, they're persisting with their efforts to invade the privacy of people who have nothing to do with this case, including by making these blatantly false allegations. We'll continue defending our users' privacy and the long-established principles of fair use," a spokesperson for OpenAI told Benzinga. The publishers also contend that testimony from an OpenAI employee contradicts the company's earlier representations during the discovery process. Specifically, they point to testimony indicating OpenAI had searched its systems for the publishers' content after previously telling the court it lacked the technical capability to conduct such searches. OpenAI has previously argued that producing ChatGPT conversation records could compromise user privacy. The publishers, however, characterize their sanctions request as a response to what they describe as missing or concealed evidence that is central to the case. As part of the sanctions motion, the publishers are also seeking attorney fees tied to what they say were unnecessary efforts to obtain evidence that should have been produced during discovery. The lawsuit adds to a growing wave of legal challenges from publishers, authors and other content creators, as courts weigh how copyright law applies to AI models trained on publicly available and copyrighted material. In March, Grammarly faced a lawsuit over the alleged use of its Expert Review AI tool. Julia Angwin, a contributing opinion editor at The New York Times, alleges that the tool used her name and others' without prior consent. This content was partially produced with the help of AI tools and was reviewed and published by Benzinga editors. Market News and Data brought to you by Benzinga APIs To add Benzinga News as your preferred source on Google, click here.
[15]
OpenAI Accused of Lying to Court in Newspaper Lawsuit | PYMNTS.com
The papers, The New York Times (NYT) and New York Daily News among them, accuse OpenAI of lying to a federal court about its ability to search its systems for evidence that it misused millions of their news stories to train its AI model, Reuters reported Thursday (July 9). "For over two years, OpenAI lied to The Times, The Daily News Plaintiffs, the public, and the court," the New York Times' lead attorney Ian Crosby said in a statement to Reuters. "It claimed searching ChatGPT outputs for copies of The Times' and the Daily News Plaintiffs' content was infeasible, burdensome, and invasive of users' privacy - while at the same time concealing that it had already done such searches." The filing also alleges that OpenAI deleted billions of relevant ChatGPT conversations or made them unsearchable. The plaintiffs are seeking sanctions including attorneys fees, as well as a finding that the company's chat logs showed OpenAI misused the papers' copyrighted material. Reached for comment by PYMNTS, a spokesperson for OpenAI shared a statement rejecting the plaintiffs' allegations. "As the Times' case weakens and they've been forced to drop claims against us, they're persisting with their efforts to invade the privacy of people who have nothing to do with this case, including by making these blatantly false allegations," the company said. "We'll continue defending our users' privacy and the long-established principles of fair use." The suit was filed by the NYT in 2023, accusing OpenAI and its benefactor Microsoft of using millions of its articles without consent to train OpenAI's ChatGPT. The NYT has since dropped some of its claims, per a recent Bloomberg Law report. As Reuters noted, this is one of many suits filed by copyright owners against AI companies for allegedly misusing their books, songs and news articles to train their systems. So far, PYMNTS wrote last month, the judicial record on cases like these shows a divide. For example, U.S. District Judge William Alsup in San Francisco called AI training "quintessentially transformative" and said copyright law "seeks to advance original works of authorship, not to protect authors against competition." But U.S. District Judge Vince Chhabria, also in San Francisco, came to a different conclusion two days later, warning that widespread AI training could impede the economic incentives that drive human creative work. Daryl Lim, H. Laddie Montague Jr. Chair in Law at Penn State Dickinson Law School, told PYMNTS that only a handful of companies can train frontier models at scale, since these firms control compute, data, cloud infrastructure, and distribution all at once. "When you train frontier models, you need to ingest vast repositories of works that may include copyrighted works," Lim said.
[16]
OpenAI faces sanctions bid as copyright case escalates
We have largely come to accept a modern digital bargain: When we interact with AI chatbots, our prompts become fodder for their ongoing training. These models feast on a constant stream of user data to survive. Now, however, that massive accumulation of chat history is exactly what has landed OpenAI in hot water. The New York Times and more than a dozen other publishers asked a federal judge on Thursday to sanction OpenAI over how the company handled evidence in their copyright lawsuit. The original case asked whether OpenAI trained ChatGPT on stolen journalism. This motion asks something narrower and more damaging: whether OpenAI lied about what it could already prove. OpenAI accused of obstruction The filing, submitted in the Southern District of New York, accuses OpenAI of a "deliberate and systemic effort to obstruct discovery," according to a Bloomberg Law report. For two years, OpenAI told the court that searching ChatGPT training data and logs for copyrighted material was not technically feasible, a Reuters report noted. Plaintiffs say that was false. The claim rests on a February deposition of Vinnie Monaco, OpenAI's privacy engineering lead, a TechCrunch report confirmed. Monaco reportedly testified that OpenAI had already searched its training corpus and built a database of roughly 78 million de-identified ChatGPT conversations before the Times filed suit in 2023. He also said it had developed an internal tool called Project Giraffe to detect reproduced text. That timeline undercuts OpenAI's central excuse, since the tools existed before the lawsuit, not because of it. boonchai wedmakawand / Getty Images OpenAI court case: disputed evidence and deleted logs Plaintiffs originally sought a sample of 120 million chat logs. OpenAI negotiated that down to 20 million, then submitted a version in December so heavily redacted that the court called it unusable, according to TechCrunch. The newspapers also allege that OpenAI deleted billions of ChatGPT conversations after a court preservation order took effect. The remedy plaintiffs want is where the real leverage lies. They are asking the judge to bar OpenAI from using its own reduced log sample as evidence and to simply rule as fact that the logs would show OpenAI reproduced their journalism, Reuters noted. Ian Crosby, the Times' lead attorney at Susman Godfrey, said OpenAI concealed what it knew for more than two years, and that resolving discovery this way would settle the case's most contested technical question without a trial. OpenAI's defense and the financial stakes OpenAI rejects that framing entirely. Spokesperson Drew Pusateri said in a statement that The Times is using privacy claims to compensate for a weakening case, and that the company will keep defending user privacy and fair use. Notably, the sanctions motion does not target Microsoft, OpenAI's co-defendant and largest financial backer, according to The Times' own report on its filing. The dollar figures already on the table elsewhere in AI litigation show what is at stake. Anthropic agreed to pay authors $1.5 billion to settle a separate training-data case, the largest AI copyright settlement to date, according to The Associated Press. That amount is a small fraction of Anthropic's $965 billion valuation, but it set a price for what a court finding against an AI company can cost. The Times itself has spent more than $28 million fighting AI companies in court, including $4.2 million in this year's first quarter alone, a Variety report confirmed. NYT shares traded near $73 this week with little visible reaction to the filing. Discovery motions rarely move markets the way settlements or verdicts do, even when the allegations are this severe. The bigger stakes as OpenAI eyes a public listing What the sanctions fight really tests is how AI companies answer for their conduct during litigation, separate from what they did while building their models. OpenAI is preparing for a public listing that bankers have discussed at valuations approaching $1 trillion, and its eventual prospectus will need to disclose litigation risk like this to new shareholders. A finding of discovery misconduct would hand publishers, and every publisher watching this case, leverage that a fair use defense alone cannot buy back. The Arena Media Brands, LLC THESTREET is a registered trademark of TheStreet, Inc. This story was originally published July 10, 2026 at 5:03 PM.
[17]
OpenAI Celebrations Around Powerful New AI Model Dampened by Fresh Media Lawsuit
The lawsuit seeks legal sanctions against OpenAI for withholding evidence and asks for transparency in how they create and train AI models OpenAI CEO Sam Altman would be wondering what he has done wrong with media houses to incur their continued wrath for such a long time. Given that he seldom acknowledges the value of copyright, the AI big boss might feel hard-done by a bunch of 15 publishers who're now seeking legal sanctions against his company for allegedly withholding evidence. OpenAI finally released its most powerful model yet - the GPT-5.6 Sol - after a few weeks delay when the Trump administration stopped the rollout over cybersecurity concerns. The company also unveiled a new tool ChatGPT Work to match Claude Cowork by Anthropic. And then the publishers struck with a lawsuit that claims OpenAI had evidence that would've helped decide a few older cases. With this filing, the publishers also sought strict legal sanctions against OpenAI, accusing the company of violating court rules and acting in bad faith during the litigation's fact discovery phase. In the latest sanctions filing, the plaintiffs said OpenAI refused to provide information showing how the company's AI systems are trained and used. "The evidence is in OpenAI's training data sets and ChatGPT output logs," the publishers said in their motion. "But instead of just producing that evidence at the start of the case and focusing on the merits of its fair use defence, OpenAI chose obstruction." The NYT had sued OpenAI back in 2023, accusing it of infringing on copyrights by using the publication's content to train ChatGPT and other technologies. In the months that followed, other publishers sued Sam Altman's AI startup, making similar accusations. Given the similarities between these case, several of them were consolidated last year. As is their wont, OpenAI put on a brave face and sought to counter the publishers. OpenAI spokesperson Drew Pusateri said the Times' case was weakening and therefore are forced to drop claims. But they're persisting with efforts to invade the privacy of people who have nothing to do with the case. Ever since the latest AI models hit the market, publishers have questioned their penchant for ignoring the IP of content creators while using it to train models. Most AI labs are facing similar charges of using copyrighted material to train their AI without having to pay for it. Some of these publishers later got themselves monetary deals with OpenAI and others. The latest motion against OpenAI came after the latter's employee had deposed. It was largely redacted in public court documents, but still indicates that OpenAI could've provided the data that the complainants sought. "For two years, OpenAI has been making misrepresentations to the court regarding its ability to search for Daily News content in its training data sets and output logs," said Steven Lieberman, counsel for The Daily News and other newspapers. Coming back to their AI model launch, GPT-5.6 has three variants - Sol the Workhorse, Terra the Intermediate and Luna the Budget-friendly that can help users expand work into several fiends such as enterprise, coding and scientific research. Sam Altman said his company's new models are more efficient and cost-effective than previous versions with Sol being 54% more token efficient in AI coding tasks. The GPT-5.6 was also described as their "strongest cybersecurity model yet, achieving frontier performance with significantly fewer tokens." The company also released a new tool called ChatGPT Work, which is designed as a workplace companion for enterprise teams, running on desktop, web, and mobile. One could call it a competitor to Claude Cowork - both helping customers with daily clerical tasks, like drafting documents, spreadsheets, and presentations. Another interesting aside to the launch was OpenAI's attempt to counter recent reports around Microsoft seeking options in lieu of OpenAI software via its own in-house models. In a blog post, the company said GPT-5.6 would support Microsoft users across the company's range of productivity apps including Word, Excel, PowerPoint and Cowork. "Our partnership with Microsoft has always been about bringing the benefits of advanced AI to more individuals and organizations, and we're excited to continue building on that shared commitment," the post said. In fact, OpenAI said it would become the "preferred model" powering Microsoft's 365 Copilot. The problem is that the company does not really clarify what the preferred status brings with it or how it would respond to Microsoft's needs in the future. Maybe, it is just Sam Altman's effort to show the world was all is hunky-dory between Microsoft and OpenAI.
[18]
Why The New York Times wants OpenAI sanctioned by court
You can access the court document from here. A group of news publishers led by The New York Times has asked a US federal court to sanction OpenAI, accusing the company of misleading the court about its ability to search for copyrighted news content used to train its AI models and of deleting or making billions of ChatGPT conversation logs unsearchable during the ongoing copyright dispute. In a sanctions motion filed in the US District Court for the Southern District of New York on July 9, the publishers alleged that OpenAI falsely told the court it could not search its training datasets or ChatGPT output logs for the publishers' copyrighted content, even though OpenAI had allegedly searched those systems for the publishers' content before the publishers sued the company. Publishers allege OpenAI hid search capabilities: "For over two years, OpenAI lied to The Times, The Daily News Plaintiffs, the public, and the court," The New York Times' lead attorney Ian Crosby said in a statement. "It claimed searching ChatGPT outputs for copies of The Times' and the Daily News Plaintiffs' content was infeasible, burdensome, and invasive of users' privacy, while at the same time concealing that it had already done such searches." The publishers also accused OpenAI of failing to preserve evidence by allegedly deleting or compressing billions of ChatGPT conversations, making them unavailable for discovery. According to the court filing, OpenAI continued deleting logs even after the court ordered it to preserve and segregate relevant ChatGPT output data. What the publishers are asking the court to do: The filing further alleges that OpenAI misrepresented its technical capabilities throughout the discovery process. It says the company repeatedly told the court it lacked tools to search its training datasets and output logs. Still, a later deposition by an OpenAI corporate witness revealed that the company had already built tools to search training data, create searchable datasets of de-identified ChatGPT conversations, and search news publishers' content. The publishers have asked the court to impose sanctions under US civil procedure rules. Among the remedies sought are a ban on OpenAI relying on a 20-million-conversation ChatGPT sample during the case, a court finding that ChatGPT logs contain or would have shown "substantial and systematic" reproduction of the publishers' copyrighted works, attorneys' fees and litigation costs, and jury instructions reflecting OpenAI's alleged destruction of evidence. The motion states: "This is a case about copying. There is no question that it happened. Nor should there be one about what was copied, how often, or to what end." It further alleges that OpenAI "chose obstruction" by refusing to produce relevant evidence and concealing its ability to search its own systems, resulting in years of unnecessary discovery disputes. The publishers also argue that OpenAI used their journalism to train its AI models and reproduced their copyrighted content in ChatGPT responses. OpenAI rejected the allegations: "As the Times' case weakens and they've been forced to drop claims against us, they're persisting with their efforts to invade the privacy of people who have nothing to do with this case, including by making these blatantly false allegations," OpenAI spokesperson Drew Pusateri said. Background: The New York Times sued OpenAI and Microsoft in December 2023, alleging that the companies used millions of its copyrighted articles without permission to train the large language models behind ChatGPT. The case later expanded to include several other news organisations, including the New York Daily News, and forms part of a broader wave of copyright lawsuits filed by publishers, authors, artists and music companies against AI developers over the use of copyrighted material for AI training.
[19]
News outlets urge a judge to sanction OpenAI in a high-stakes AI copyright fight
NEW YORK -- The New York Times, the Daily News and other media outlets are asking a federal judge to impose sanctions on OpenAI, escalating a fight over artificial intelligence and copyright that could shape the future of a struggling news industry. The newspapers allege the ChatGPT maker is hiding evidence important to what could be a landmark copyright infringement trial over how OpenAI and its business partner, Microsoft, built their AI technologies using millions of news articles. At issue is whether AI chatbots are unfairly competing as an information source, siphoning off web traffic without doing the journalistic work involved in gathering the news. A filing Thursday in a Manhattan federal courthouse alleges OpenAI "chose obstruction" over releasing datasets and ChatGPT logs that could show how the AI system used copyrighted news content. The plaintiffs are asking the judge to penalize the company for "discovery misconduct" that could distort evidence, saying a recent deposition of an OpenAI employee contradicts the company's earlier claims. New York Daily News attorney Steven Lieberman said OpenAI has been "making misrepresentations" for two years about its ability to search for copyrighted content in its AI training datasets and logs. "This motion asks the court to punish OpenAI for hiding and destroying evidence showing how ChatGPT was trained on stolen journalism," said Lieberman, who represents the Daily News and seven of its sister papers. OpenAI didn't immediately respond to a request for comment Thursday. The New York Times sued OpenAI and Microsoft in late 2023, about a year after ChatGPT's debut sparked a commercial AI boom and began changing the way people search for information online. The threat to news publications became even more apparent when Google in 2024 introduced AI-generated summaries at the top of online search results, cutting off the advertising dollars that come when people click a link to the information's original source. The Times has since been joined by other news organizations, including MediaNews Group-owned newspapers the Daily News and the Chicago Tribune, digital media publisher Ziff Davis and the nonprofit Center for Investigative Reporting. OpenAI and other tech companies have argued the process of training their AI systems on digitized books, online articles and other writings found on the internet is protected by the "fair use" doctrine of U.S. copyright law. It's a theory being tested in dozens of lawsuits as visual artists, novelists, music record labels and other creative industries take AI companies to court, with mixed results. In the case involving the biggest copyright settlement so far, OpenAI rival Anthropic agreed to pay book authors US$1.5 billion for training its chatbot Claude on their pirated works -- an amount that represents a small fraction of Anthropic's US$965 billion market valuation as it prepares to become publicly traded. The New York Times' arguments are different from those brought by book authors. In its original lawsuit and an amended complaint filed last month, it focused on the unfair competition of companies that "seek to free-ride on The Times's massive investment in its journalism by using it to build substitutive products without permission or payment." The Times has already spent more than US$28 million on fighting AI companies in court, according to filings with financial regulators that disclose its litigation costs. The costs include another lawsuit the newspaper filed last year against AI company Perplexity. Among the sanctions sought by the newspapers Thursday are attorney fees that would pay for the efforts to secure "improperly withheld" evidence. The mounting costs come as a growing number of media organizations have signed licensing deals with OpenAI and other AI companies such as Google and Facebook parent Meta that typically pay the outlet a fee to be able to train AI systems on their news feeds or archives. The Associated Press was the first to announce such a deal with OpenAI in 2023. ___ Matt O'brien And Jocelyn Noveck, The Associated Press
[20]
OpenAI accused of misleading court over ChatGPT data in lawsuit filed by publishers: Here is what we know
The newspapers have asked the court to impose sanctions on OpenAI. OpenAI is facing new allegations in its ongoing copyright lawsuit with a group of newspapers led by The New York Times and the New York Daily News. In a new court filing, the outlets have accused the AI company of misleading the court about its ability to search its AI systems for their copyrighted content. The filing was submitted to a federal court in Manhattan on Thursday, reports Reuters. OpenAI had told the court it could not search its large language models for copies of their articles. The outlets argue that this statement was false because the company had already carried out such searches before the first lawsuit was filed. The newspapers have asked the court to impose sanctions on OpenAI. They are seeking attorneys' fees and also want the court to rule that OpenAI's ChatGPT records show the company used their copyrighted news articles without permission. Also read: OpenAI introduces ChatGPT Work, an AI agent powered by GPT 5.6 and Codex: Here is what it can do The filing also accuses OpenAI of deleting billions of relevant ChatGPT conversations or making them impossible to search. OpenAI has previously argued that sharing ChatGPT conversation logs could put users' privacy at risk. Responding to the latest filing, OpenAI spokesperson Drew Pusateri rejected the allegations. He said, "As the Times' case weakens and they've been forced to drop claims against us, they're persisting with their efforts to invade the privacy of people who have nothing to do with this case, including by making these blatantly false allegations." The legal battle began in 2023 when The New York Times sued OpenAI and Microsoft. The lawsuit claims that the companies used millions of newspaper articles without permission to train the AI model that powers ChatGPT. Last month, however, The New York Times dropped one of its secondary copyright infringement claims in an updated complaint. Also read: WhatsApp username row: Meta submits reply to notice, response under examination "For over two years, OpenAI lied to The Times, The Daily News Plaintiffs, the public, and the court," the New York Times' lead attorney Ian Crosby was quoted as saying in the report. "It claimed searching ChatGPT outputs for copies of The Times' and the Daily News Plaintiffs' content was infeasible, burdensome, and invasive of users' privacy - while at the same time concealing that it had already done such searches." New York Daily News attorney Steven Lieberman said the motion "asks the court to punish OpenAI for hiding and destroying evidence showing how ChatGPT was trained on stolen journalism." The case is one of several lawsuits filed by copyright owners against AI companies. Authors, artists and music labels have also accused companies such as OpenAI, Anthropic and Meta of using copyrighted material without permission to train AI models.
Share
Copy Link
The New York Times and 16 other publishers are asking a federal judge to sanction OpenAI for allegedly concealing evidence in their copyright infringement case. Court filings claim OpenAI lied for two years about its ability to search training data and ChatGPT logs, while secretly maintaining a database of 78 million conversations that could prove the AI company used copyrighted journalism without permission.
The New York Times and a coalition of 16 other publishers filed a sanctions motion Thursday in Manhattan federal court, accusing OpenAI of systematically hiding evidence and lying about its technical capabilities in what could become a landmark copyright infringement case
1
. The allegations center on OpenAI's repeated claims that it lacked the ability to search its systems for copyrighted content, statements now revealed to be false according to testimony from OpenAI privacy engineer Vincent Monaco during an April court-ordered deposition2
.The news outlets are requesting legal sanctions including financial penalties and attorney fees, arguing that OpenAI and Microsoft built their AI technologies using millions of news articles without permission
3
. "For over two years, OpenAI lied to The Times, The Daily News Plaintiffs, the public, and the court," said Ian Crosby, lead counsel for the plaintiffs, in a statement4
.
Source: MediaNama
Among the most damaging revelations in the ChatGPT copyright trial, Monaco's testimony allegedly disclosed that OpenAI had already amassed a database of approximately 78 million de-identified ChatGPT conversations before The New York Times even filed its lawsuit in 2023
2
. The company was using this data internally to evaluate how much it was infringing on copyrighted works, yet never disclosed its existence to the court or plaintiffs during two years of discovery proceedings1
.Additionally, OpenAI allegedly maintained a separate sample of 10 million logs that had already been de-identified and could have been made available early in the discovery process
1
. The publishers claim OpenAI had already searched these samples for New York Times content as part of research into "creating a filter that could be used to block the regurgitation of copyrighted content" through an initiative called "Project Giraffe"2
.
Source: PYMNTS
News outlets urge judge to sanction OpenAI for what they characterize as systematic obstruction during the discovery process. Publishers originally requested a sample of 120 million chat logs, but OpenAI negotiated this down to just 20 million, citing technical burdens and user privacy concerns
2
. When OpenAI finally submitted this smaller sample in December, it had made 19 billion redactions using AI, rendering the evidence "unusable" according to the court1
.The plaintiffs further allege that OpenAI deleted billions of relevant ChatGPT conversations after the lawsuit was filed, potentially violating the court's preservation order, and substituted millions of logs in the requested sample
2
. "This motion asks the court to punish OpenAI for hiding and destroying evidence showing how ChatGPT was trained on stolen journalism," said New York Daily News attorney Steven Lieberman5
.Related Stories
The OpenAI withholding evidence in copyright lawsuits allegations strike at the heart of how large language models are trained and whether this constitutes fair use under intellectual property law. OpenAI and Microsoft have consistently argued that training AI systems on copyrighted material falls under fair use doctrine, while publishers contend the companies are building "substitutive products" that compete directly with original journalism
5
.The journalism industry faces mounting pressure from AI-generated search summaries that reduce traffic to original sources. Some data shows small publishers have experienced a 60% traffic drop, with projections suggesting declines exceeding 40% by 2029
3
. The New York Times has already spent more than $28 million fighting AI companies in court, according to financial regulatory filings5
.
Source: BNN
The New York Times-led group asks court to sanction OpenAI by preventing the company from using the compromised 20 million log sample as evidence, accepting as fact that chat logs would have shown substantial regurgitation of copyrighted content, and requiring OpenAI to pay legal fees incurred while pursuing improperly withheld evidence
2
.OpenAI spokesperson Drew Pusateri denied the allegations, stating: "As the Times' case weakens and they've been forced to drop claims against us, they're persisting with their efforts to invade the privacy of people who have nothing to do with this case, including by making these blatantly false allegations"
2
. However, NYT spokesperson Graham James disputed that dropping one claim weakened their case, noting they added claims against Microsoft while maintaining their core allegations .The outcome could establish critical precedents for how AI companies access and use training data, potentially reshaping the relationship between tech giants and content creators across multiple industries facing similar legal challenges.
Summarized by
Navi
08 Nov 2024•Policy and Regulation

30 Nov 2024•Policy and Regulation

25 Jun 2026•Policy and Regulation
