WikiHow Sues OpenAI for Copyright Infringement Over Unauthorized AI Training on 11,000+ Articles

2 Sources

Share

WikiHow filed a federal lawsuit against OpenAI, alleging the AI company scraped over 11,000 copyrighted how-to articles without permission to train ChatGPT. The suit claims OpenAI's mass-scale copying threatens WikiHow's revenue by reproducing its content and depriving it of web traffic.

WikiHow Sues OpenAI Over Unauthorized Use of Copyrighted Content

WikiHow sues OpenAI in a Manhattan federal court lawsuit filed on August 24, alleging copyright infringement through unauthorized AI training on its instructional content. The do-it-yourself platform claims OpenAI scraped over 11,000 of its how-to articles without permission to train ChatGPT and other AI models.

1

The lawsuit alleges OpenAI infringed more than 1,200 of WikiHow's registered copyrights, with the platform seeking an unspecified amount of monetary damages and a court order blocking further infringement.

Mass-Scale Copying and Scraping Operations

The lawsuit details how OpenAI allegedly engaged in mass-scale copying of WikiHow's copyrighted articles at multiple stages of the AI development process. WikiHow has documented extensive scraping activity, with OpenAI's crawlers reaching its articles more than 185,000 times between March and May 2025, and 148,529 times between May and July 2026.

2

Combined with Common Crawl and OpenAI's bots, WikiHow's websites were scraped over 300,000 times since May 2025. The platform alleges OpenAI uses obfuscated user agents and IP addresses, suggesting actual scraping frequency may be even higher than documented.

Training Data Sources and GPT Model Development

WikiHow identifies specific datasets used to train ChatGPT that allegedly contain its misused instructional material. OpenAI built the "WebText" corpus for GPT-2 by scraping high-quality websites, including WikiHow.

2

The company similarly created "WebText2" for GPT-3 training, again targeting quality sites like WikiHow. The lawsuit alleges OpenAI used Common Crawl, WebText, WebText2, or similar datasets to train GPT-4 and subsequent models, incorporating at least 11,211 of WikiHow's articles without authorization. While OpenAI has not disclosed the contents of datasets used for GPT-4 or later models, the company previously published information about training early-generation models.

Retrieval-Augmented Generation Systems Expand Copyright Violations

Beyond initial training, WikiHow accuses OpenAI of using retrieval-augmented generation (RAG) systems that retrieve, copy, and use its copyrighted articles to expand ChatGPT's knowledge base.

2

The lawsuit presents evidence where GPT-4 identified a copyrighted image from WikiHow's article "How to Restring Your Guitar at Home: Step-by-Step Guide." When questioned about this identification capability, GPT-4 reportedly stated it compares filenames against other WikiHow images and that its training data provided "enough exposure to WikiHow content to be able to generalize their visual patterns and formats effectively."

Source: MediaNama

Source: MediaNama

Economic Impact and Market Displacement Concerns

WikiHow argues that ChatGPT now reproduces and repackages its instructions, depriving WikiHow of web traffic and threatening its advertising revenue model. "Having ingested wikiHow's articles, those models now produce competing how-to content on the same subjects, at a fraction of the time, effort, and cost of researching, writing, and editing a wikiHow article," the lawsuit states.

1

The platform claims this substitution cuts its revenue and undermines its incentive to continue producing quality articles. ChatGPT allegedly reproduces WikiHow's text in response to user prompts, allowing readers to obtain the substance of WikiHow articles without visiting the site.

OpenAI's Fair Use Defense and Industry Context

OpenAI responded to the lawsuit by asserting its AI models are "trained on publicly available data and grounded in fair use."

1

This defense mirrors OpenAI's position in other legal challenges. The WikiHow lawsuit joins a wave of cases brought by publishers, authors, and copyright holders against AI companies over alleged misuse of copyrighted material to train chatbots like ChatGPT, Anthropic's Claude, Meta's Llama, and Google's Gemini. Earlier this year, Encyclopedia Britannica and Merriam-Webster sued OpenAI with similar allegations, claiming GPT-4 "memorized" their content and generated near-verbatim outputs. The New York Times has also filed an ongoing lawsuit against OpenAI, accusing the company of copying massive amounts of copyrighted content.

2

What This Means for Content Creators and AI Development

The lawsuit highlights growing tensions between AI companies and content creators over training data practices. WikiHow emphasizes that it "has paid writers and editors to produce the internet's largest library of human-authored, expert-verified how-to instructions," only to have OpenAI allegedly take that library "without a license and without payment."

2

The case raises questions about whether AI companies can continue relying on fair use arguments when scraping copyrighted articles at scale. Watch for how courts balance innovation in AI development against copyright protections, as these decisions will shape future AI training practices and potentially require companies to negotiate licensing agreements with content creators.

Today's Top Stories

© 2026 TheOutpost.AI All rights reserved