2 Sources
[1]
WikiHow sues OpenAI for copyright infringement over AI training
Aug 24 (Reuters) - Do-it-yourself website WikiHow sued OpenAI in Manhattan federal court for allegedly misusing its instructional material to train its popular chatbot ChatGPT. WikiHow's lawsuit, opens new tab, filed on Friday, said that the AI company scraped thousands of its how-to articles to help train ChatGPT to respond to human prompts. The lawsuit is part of a wave of cases brought by publishers, authors and other copyright holders against AI companies over the alleged misuse of their material to train chatbots like ChatGPT, Anthropic's Claude, Meta's â Llama and Google's Gemini. An OpenAI spokesperson said on Monday that the company's AI models are "trained on publicly available data and grounded in fair use." Spokespeople and attorneys for wikiHow did not immediately respond to a request for comment on the new lawsuit on Monday. WikiHow is a crowdsourced platform that hosts instructions for DIY projects. The lawsuit said that OpenAI copied more than 11,000 of its articles without permission to train its GPT large language models. WikiHow also said that â ChatGPT reproduces its text in response to user prompts and threatens to displace its market. "Having ingested wikiHow's articles, those models now produce competing how-to content on the same subjects, at a fraction of the time, effort, and cost of researching, writing, and â editing a wikiHow article," the lawsuit said. "The substitution cuts wikiHow's revenue and, over time, its reason to keep producing the articles at all." The lawsuit alleges OpenAI infringed more than â 1,200 of wikiHow's copyrights. WikiHow requested an unspecified amount of monetary damages and a court order blocking OpenAI from infringing its copyrights. The case is WikiHow â Inc v. OpenAI Inc, U.S. District Court for the Southern District of New York, No. 1:26-cv-07171. For WikiHow: Rohit Nath and Alejandra Salinas of Susman Godfrey For OpenAI: attorney information not yet available Reporting by Blake Brittain in Washington Our Standards: The Thomson Reuters Trust Principles., opens new tab * Suggested Topics: * Legal Industry * Corporate Counsel * Intellectual Property Blake Brittain Thomson Reuters Blake Brittain reports on intellectual property law, including patents, trademarks, copyrights and trade secrets, for Reuters Legal. He has previously written for Bloomberg Law and Thomson Reuters Practical Law and practiced as an attorney.
[2]
WikiHow sues OpenAI over AI training on copyrighted articles
WikiHow has filed a lawsuit against OpenAI, alleging that OpenAI scraped, copied and ingested its copyrighted content without permission or compensation to train its AI models. "WikiHow has paid writers and editors to produce the internet's largest library of human-authored, expert-verified how-to instructions. Defendants took that library, without a license and without payment, and used it to build and operate ChatGPT -- a product that now supplies the substance of WikiHow's articles to readers who never arrive at wikiHow's pages," the lawsuit says. OpenAI allegedly infringed on WikiHow's copyrighted work on a large scale. WikiHow says that OpenAI infringes on its copyrighted works "at every stage of the generative AI process behind ChatGPT" -- while building training data, retrieving articles at runtime and when ChatGPT delivers output to users. WikiHow identified at least three ways in which it alleges OpenAI violated its exclusive rights under the Copyright Act: * OpenAI engages in "mass-scale copying" of WikiHow's copyrighted articles to train its LLMs. * Separately, OpenAI allegedly retrieves, copies and uses WikiHow's copyrighted articles through retrieval-augmented generation (RAG) systems that expand the existing knowledge base of its LLMs. * WikiHow has also accused OpenAI of cannibalizing its web traffic. "ChatGPT reproduces and repackages WikiHow's instructions in its own answers, so that a reader obtains the substance of a WikiHow article without ever visiting wikiHow. That deprives wikiHow of the page visits on which its advertising revenue depends." The scale of infringement. "Every copy OpenAI made is a separate act of infringement, and the copying has not stopped," said WikiHow. * OpenAI's crawlers allegedly reached WikiHow's articles more than 185,000 times between March and May 2025 and 148,529 times between May and July 2026. * Common Crawl and OpenAI's bots have scraped WikiHow's websites more than 300,000 times since May 2025. How does OpenAI train its AI models? According to the lawsuit, OpenAI has not disclosed the contents of the datasets used to train GPT-4 or subsequent models. The copying that occurs within its training pipelines and RAG systems is not visible from outside those systems. * However, the company was not always secretive about its training data. It published some information about how it trained its early-generation LLMs. GPT-2's training data included an internal corpus OpenAI built and named "WebText," assembled by scraping the web with an emphasis on "document quality". That corpus includes thousands of pages from wikiHow's website. * Likewise, OpenAI built the "WebText2" corpus used to train GPT-3 by crawling and scraping the web, again favoring high-quality sites such as WikiHow's. GPT-3's most heavily weighted training input was a dataset called "Common Crawl." OpenAI allegedly used Common Crawl to train GPT-4. The lawsuit alleges that OpenAI used Common Crawl, WebText, WebText2, or similar datasets to train GPT-4 and subsequent models, and those datasets included at least 11,211 of WikiHow's articles. * WikiHow cited its website logs as further evidence of OpenAI's crawling, scraping, and copying of its articles. WikiHow's website allegedly received more than 185,000 visits from OpenAI's published IP addresses between March 27, 2025 and May 16, 2025. * According to the lawsuit, OpenAI also uses obfuscated user agents and IP addresses, so those figures understate the frequency of their crawling, scraping, and copying. * The lawsuit also includes a copyrighted image, which, WikiHow claimed, GPT-4 readily identifies as coming from the WikiHow article "How to Restring Your Guitar at Home: Step-by-Step Guide." * When asked how that identification was possible, GPT-4 reportedly said that it compares the file name against the filenames of other wikiHow images and that its training data gave it "enough exposure to WikiHow content to be able to generalize their visual patterns and formats effectively." * "The model's reference to its own exposure to WikiHow content supports the inference that WikiHow's articles, including the works asserted here, were copied into the data used to train GPT-4 and later models," the lawsuit claims. * Overall, WikiHow alleged that OpenAI has infringed on at least 1,200 of its registered copyrighted works to date. It has asked the court to block OpenAI from continuing to store or use its works and sought an undisclosed amount of monetary damages. Not the first copyright lawsuit against an AI company. Earlier this year, Encyclopedia Britannica and dictionary publisher Merriam-Webster sued OpenAI, alleging it used their copyrighted content to train its AI. According to the lawsuit, GPT-4 "memorized" their content and generated near-verbatim outputs on demand. The New York Times has made similar claims in its ongoing lawsuit against OpenAI, including accusing the AI company of copying massive amounts of its copyrighted content. Meanwhile, CNN has sued Perplexity for alleged copyright and trademark infringement. Why did the Delhi HC deny ANI interim relief in its copyright case against OpenAI? Last month, the Delhi High Court denied an interim injunction to news agency ANI's claim of copyright infringement against OpenAI. * The court observed that "OpenAI's act of storing ANI's works does not amount to copyright infringement...since outputs were not similar to ANI's outputs." * All the ANI articles cited in support of the claim were published in August or September 2024, after the training cut-off dates of GPT-4 (April 2022) and GPT-4o (April 2024). Accordingly, those specific articles could not have been part of the training data. Therefore, the claim of memorization and regurgitation of those works could not be sustained.
Share
Copy Link
WikiHow filed a federal lawsuit against OpenAI, alleging the AI company scraped over 11,000 copyrighted how-to articles without permission to train ChatGPT. The suit claims OpenAI's mass-scale copying threatens WikiHow's revenue by reproducing its content and depriving it of web traffic.
WikiHow sues OpenAI in a Manhattan federal court lawsuit filed on August 24, alleging copyright infringement through unauthorized AI training on its instructional content. The do-it-yourself platform claims OpenAI scraped over 11,000 of its how-to articles without permission to train ChatGPT and other AI models.
1
The lawsuit alleges OpenAI infringed more than 1,200 of WikiHow's registered copyrights, with the platform seeking an unspecified amount of monetary damages and a court order blocking further infringement.The lawsuit details how OpenAI allegedly engaged in mass-scale copying of WikiHow's copyrighted articles at multiple stages of the AI development process. WikiHow has documented extensive scraping activity, with OpenAI's crawlers reaching its articles more than 185,000 times between March and May 2025, and 148,529 times between May and July 2026.
2
Combined with Common Crawl and OpenAI's bots, WikiHow's websites were scraped over 300,000 times since May 2025. The platform alleges OpenAI uses obfuscated user agents and IP addresses, suggesting actual scraping frequency may be even higher than documented.WikiHow identifies specific datasets used to train ChatGPT that allegedly contain its misused instructional material. OpenAI built the "WebText" corpus for GPT-2 by scraping high-quality websites, including WikiHow.
2
The company similarly created "WebText2" for GPT-3 training, again targeting quality sites like WikiHow. The lawsuit alleges OpenAI used Common Crawl, WebText, WebText2, or similar datasets to train GPT-4 and subsequent models, incorporating at least 11,211 of WikiHow's articles without authorization. While OpenAI has not disclosed the contents of datasets used for GPT-4 or later models, the company previously published information about training early-generation models.Beyond initial training, WikiHow accuses OpenAI of using retrieval-augmented generation (RAG) systems that retrieve, copy, and use its copyrighted articles to expand ChatGPT's knowledge base.
2
The lawsuit presents evidence where GPT-4 identified a copyrighted image from WikiHow's article "How to Restring Your Guitar at Home: Step-by-Step Guide." When questioned about this identification capability, GPT-4 reportedly stated it compares filenames against other WikiHow images and that its training data provided "enough exposure to WikiHow content to be able to generalize their visual patterns and formats effectively."
Source: MediaNama
WikiHow argues that ChatGPT now reproduces and repackages its instructions, depriving WikiHow of web traffic and threatening its advertising revenue model. "Having ingested wikiHow's articles, those models now produce competing how-to content on the same subjects, at a fraction of the time, effort, and cost of researching, writing, and editing a wikiHow article," the lawsuit states.
1
The platform claims this substitution cuts its revenue and undermines its incentive to continue producing quality articles. ChatGPT allegedly reproduces WikiHow's text in response to user prompts, allowing readers to obtain the substance of WikiHow articles without visiting the site.Related Stories
OpenAI responded to the lawsuit by asserting its AI models are "trained on publicly available data and grounded in fair use."
1
This defense mirrors OpenAI's position in other legal challenges. The WikiHow lawsuit joins a wave of cases brought by publishers, authors, and copyright holders against AI companies over alleged misuse of copyrighted material to train chatbots like ChatGPT, Anthropic's Claude, Meta's Llama, and Google's Gemini. Earlier this year, Encyclopedia Britannica and Merriam-Webster sued OpenAI with similar allegations, claiming GPT-4 "memorized" their content and generated near-verbatim outputs. The New York Times has also filed an ongoing lawsuit against OpenAI, accusing the company of copying massive amounts of copyrighted content.2
The lawsuit highlights growing tensions between AI companies and content creators over training data practices. WikiHow emphasizes that it "has paid writers and editors to produce the internet's largest library of human-authored, expert-verified how-to instructions," only to have OpenAI allegedly take that library "without a license and without payment."
2
The case raises questions about whether AI companies can continue relying on fair use arguments when scraping copyrighted articles at scale. Watch for how courts balance innovation in AI development against copyright protections, as these decisions will shape future AI training practices and potentially require companies to negotiate licensing agreements with content creators.Summarized by
Navi
16 Mar 2026â¢Policy and Regulation

30 Nov 2024â¢Policy and Regulation

08 Nov 2024â¢Policy and Regulation

1
Technology

2
Policy and Regulation

3
Technology
