NX
App

WikiHow Sues OpenAI: The Expanding Battle Over AI Training Data and Copyright

x/Business - Work Smarter. Grow Faster. x/business ·
WikiHow Sues OpenAI: The Expanding Battle Over AI Training Data and Copyright

WikiHow Sues OpenAI: The Expanding Battle Over AI Training Data and Copyright

By Tony | August 25, 2026


The copyright battlefield between AI companies and content publishers just got another major combatant. WikiHow — the world's largest collaborative "how-to" guide website — filed a federal lawsuit against OpenAI on August 21, 2026, accusing the AI giant of scraping over 11,000 tutorial articles without permission to train ChatGPT and its underlying GPT large language models.


The Allegations at a Glance

WikiHow's lawsuit, filed in the U.S. District Court for the Southern District of New York (Case No. 1:2026-cv-07171), lays out a case that is becoming distressingly familiar in the AI era:

  • Mass scraping: OpenAI allegedly scraped more than 11,000 WikiHow articles covering topics from everyday tasks to specialized professional skills.
  • Copyright infringement: The suit identifies at least 1,200 registered copyrights that OpenAI is accused of violating.
  • Direct competition: WikiHow argues that ChatGPT can now reproduce WikiHow-style text and generate competing tutorial content on the same topics — at a fraction of the time, effort, and cost.

The core of WikiHow's argument is about economic displacement. As the complaint states:

"These models, after absorbing WikiHow's articles, can now produce competing tutorial content at a far lower cost in time, effort, and money than it takes to research, write, and edit WikiHow articles. This substitution reduces WikiHow's revenue — and over time, it eliminates the incentive to keep creating articles at all."


OpenAI's Response: The Fair Use Defense

OpenAI responded on August 24 through a spokesperson, stating that the company's models are "trained on publicly available data and grounded in fair use." This has been OpenAI's consistent position across multiple lawsuits — that training on publicly accessible web content constitutes fair use under U.S. copyright law.

WikiHow's legal team and spokesperson did not immediately respond to media requests for comment.


The Bigger Picture: A Wave of Publisher Lawsuits

WikiHow is far from alone. The lawsuit joins a growing roster of publishers and content creators taking OpenAI to court:

Plaintiff Filed Articles Allegedly Scraped Status
The New York Times Dec 2023 Millions of articles Active discovery; trial expected late 2026/early 2027
Authors Guild (class action) Sep 2023 Full books by Grisham, Martin, and others Active in MDL
Encyclopedia Britannica & Merriam-Webster Mar 2026 ~100,000 articles Pending in SDNY
Nine Regional Newspapers (San Diego Union-Tribune et al.) Nov 2025 Hundreds of thousands Pending in SDNY
WikiHow Aug 2026 11,000+ articles Newly filed

All of these cases have been consolidated or are being handled in the Southern District of New York, where Judge Sidney Stein is overseeing the multidistrict litigation (MDL No. 3143). In October 2025, Judge Stein denied OpenAI's motion to dismiss the consolidated direct copyright infringement claims — a significant early win for the publishers that allowed the cases to move into discovery.


The Fair Use Question: Data Scraping in the Spotlight

At the heart of the legal battle sits a question that could reshape the internet: Is training AI models on publicly available web content fair use?

OpenAI and other AI companies have leaned heavily on the argument that training constitutes "transformative use" — that the models don't reproduce content verbatim but rather learn patterns and relationships from it. But the publishers counter that:

  1. The outputs compete directly with their original content — if ChatGPT can give you a step-by-step tutorial on how to fix a leaky faucet, why visit WikiHow?
  2. Verbatim reproduction has been documented — The New York Times famously demonstrated ChatGPT outputs that mirrored its articles nearly word-for-word.
  3. Market harm is real — Wikipedia's traffic has fallen an estimated 8% as AI platforms absorb queries that once went to the encyclopedia. Google AI Overviews have similarly cannibalized publisher traffic.

The publishers' position was strengthened when the court denied OpenAI's motion to dismiss the core copyright claims. This means the case will proceed to discovery — where OpenAI may be forced to reveal exactly what data was used to train its models, how it was acquired, and the technical relationship between training inputs and model outputs.


The Economic Calculus

WikiHow's lawsuit highlights a cold economic reality. Producing high-quality, fact-checked instructional content requires real investment: subject-matter expertise, professional writing and editing, illustration, and ongoing updates. If AI models can ingest all of that work and then generate passable substitutes at near-zero marginal cost, the economic model that supports content creation breaks down.

Research indicates that AI-generated search results can cannibalize publisher traffic while converting at rates 4.4 times higher than traditional organic search — meaning the AI platforms capture the most commercially valuable queries while the content creators who made those answers possible see diminishing returns.


The Licensing Alternative

Not every publisher is suing. A parallel trend has seen major content platforms signing licensing deals with AI companies:

  • Reddit signed a licensing agreement with OpenAI specifically structured, as OpenAI Chair Bret Taylor testified, "to avoid litigation."
  • News Corp and other major publishers have struck data-licensing arrangements.
  • Some publishers have taken a hybrid approach — licensing some content while litigating over past scraping.

This two-track strategy — sue over past infringement, license for future use — is becoming the standard playbook for major content owners navigating the AI era.


What's at Stake

The WikiHow lawsuit, while smaller in scale than the NYT or Britannica cases, is significant for several reasons:

1. It represents a different category of publisher

WikiHow isn't a news organization or an encyclopedia — it's a practical, how-to knowledge base. If WikiHow's content can be freely scraped and repurposed by AI, what user-generated or collaboratively-created content is safe?

2. The competition argument is unusually direct

A ChatGPT-generated "how to tie a tie" directly substitutes for a WikiHow article on the same topic in a way that AI-generated news summaries don't perfectly substitute for original reporting. The substitution is near-perfect.

3. It tests the limits of "publicly available"

WikiHow's content is free to read — but its terms of service and copyright registrations establish clear legal boundaries. The case will help define whether "publicly accessible" means "free to take."


What Comes Next

The WikiHow case will likely be consolidated into the existing MDL proceedings before Judge Stein. Discovery is expected to continue through 2026, with a trial potentially in early 2027. Unless, of course, the parties settle — as has happened in several AI copyright cases, including the Anthropic settlement that required destruction of pirated training files and roughly $3,000 per covered work.

The courts are now drawing lines around what AI companies can and cannot do. The outcome of these cases will determine whether the internet's vast store of human-created knowledge remains a commons that supports the creators who build it — or becomes raw material for AI models that eventually render those creators obsolete.

For now, WikiHow has joined the fight. And the list of plaintiffs is only getting longer.


Sources: Reuters, IT之家, Justia Dockets, CourtListener, AI Lawsuit Tracker, TradingView News, TechCrunch, The Washington Times, Courthouse News Service

·