by ChatGPT-5.6
The lawsuit brought by wikiHow against OpenAI is superficially familiar. It joins the growing body of litigation in which authors, publishers and other rights holders argue that generative-AI developers copied copyrighted material without permission to build commercial AI systems. But the complaint is more interesting than a straightforward “you trained on our content without a licence” case.
wikiHow has constructed the case around a chain of alleged conduct: acquisition, training, retrieval, output, commercial substitution and removal of attribution information. Its central proposition is that OpenAI took professionally produced instructional content, used that content to build and operate ChatGPT, and then deployed ChatGPT as a lower-cost substitute for the very webpages from which the information was obtained. The complaint describes this as the “unlicensed taking of a publisher’s written work to build a product that competes with it.”
That distinction matters. A court may eventually be willing to regard some forms of LLM training as transformative fair use while still concluding that retrieving a publisher’s current article, reproducing substantial portions of it, stripping attribution and supplying the result to a user who would otherwise have visited the publisher is a very different activity.
wikiHow’s first grievance is the most familiar: unauthorised copying for model training. It says OpenAI incorporated thousands of wikiHow pages into datasets including WebText, WebText2 and Common Crawl-derived datasets and alleges that later GPT models also incorporated its material. wikiHow owns 1,211 copyright registrations covering 11,211 articles asserted in the case; the company argues that its copyright protects not merely facts or instructions but its particular wording, selection, ordering of steps, warnings and explanatory material.
The second grievance is more consequential: continued copying through retrieval-augmented generation or comparable retrieval systems. The complaint argues that this is not merely historical training. When a user asks ChatGPT a question, OpenAI may allegedly retrieve wikiHow material again at runtime and feed that material into the model as context. If proved, this turns part of the case from an abstract argument about what happens during model training into a much more conventional copyright question: what authority does an AI service have to fetch, copy and repurpose an identifiable copyrighted work while providing a competing answer?
Third is infringing output. wikiHow produces several examples in which GPT-4 returned passages from named wikiHow articles. The screenshots on pages 20–25 of the complaint show ChatGPT responses paired with the corresponding source passages, including examples concerning mother-in-law boundaries, wedding dreams, Aries characteristics and “soul ties.” The overlaps include substantial common wording rather than merely common facts.
Fourth, wikiHow alleges removal of copyright management information (CMI) under §1202 of the DMCA. This is not framed merely as “ChatGPT forgot to cite us.” The complaint alleges that the extraction pipeline itself removed titles, bylines, copyright notices, terms of use and other identifying material. It specifically identifies the Dragnet and Newspaper extraction tools and argues that those systems were designed to isolate article body text while discarding surrounding webpage elements.
Finally, all of these allegations feed into the commercial grievance: substitution. wikiHow is free to users and finances the creation, checking and updating of its articles through advertising and licensing. Its argument is therefore unusually simple: if someone asks ChatGPT “how do I do X?” and receives the answer there, the user may have no reason to visit the wikiHow page that financed the production of that answer. The complaint says wikiHow has already suffered reduced traffic, advertising and licensing revenues, although the complaint itself does not yet provide the econometric evidence necessary to establish how much of that reduction was caused by ChatGPT.
OpenAI’s public position, reported by Reuters, is that its models are trained on publicly available information and that this activity is protected by fair use.
The case is stronger as a pleading than it is yet as a fully proven case. That distinction is essential. A complaint contains allegations. Discovery will determine which of them can actually be demonstrated.
Several parts are comparatively strong.
The ownership position appears well developed. wikiHow has identified registered works and explains the employment and assignment arrangements through which it acquired the relevant rights. That eliminates one problem that has complicated some earlier AI cases.
The web-access evidence is also significant. wikiHow alleges more than 185,000 visits from OpenAI-published IP addresses during a period in 2025. More importantly for knowledge and willfulness, wikiHow says that it prohibited GPTBot through robots.txt from August 2023, subsequently blocked OAI-SearchBot and ChatGPT-User, notified OpenAI’s legal department directly in December 2024, and nevertheless recorded another 148,529 OpenAI-bot visits between May and July 2026. If the technical logs substantiate the attribution and show that restricted crawlers actually obtained protected content, that is considerably more compelling evidence than merely demonstrating that a work happened to exist somewhere in Common Crawl.
The output evidence is also useful because there are readily observable textual overlaps. In several screenshots, one does not need an elaborate model-inference theory to see common wording.
But there are important weaknesses.
Crawler access is not synonymous with model training. A server log can prove that a bot requested a page. It does not by itself prove that the resulting material entered a particular GPT training dataset. Some requests could relate to search, browsing, indexing, safety analysis or retrieval. Discovery of internal lineage records will therefore matter enormously.
Similarly, the complaint’s inference that later GPT models necessarily involved a “fresh ingestion” of wikiHow articles should not simply be accepted as a technical fact. Model development can involve continued pre-training, fine-tuning, synthetic data, distillation and other processes; a new model release does not inherently demonstrate that the original web corpus was re-downloaded in its entirety.
One piece of evidence should be given particularly little weight: GPT-4 reportedly said that its “training data” had exposed it sufficiently to wikiHow to recognise its visual style. A model’s statements about its own training provenance are not reliable forensic evidence of its training dataset. LLMs are not authoritative witnesses to their internal provenance. The server logs, historical WebText documentation and internal OpenAI data records are substantially more important.
There is also a weakness in the reproduction examples. The lawyers deliberately prompted ChatGPT with the name of the article and expressly asked for verbatim text. wikiHow acknowledges this and says it did so to demonstrate what the system could reproduce. That is legitimate as a stress test, but it does not establish how ChatGPT behaves in ordinary use. OpenAI can argue that the examples resemble adversarial extraction tests rather than normal consumer behaviour.
The complaint consequently needs a second body of evidence: thousands of ordinary, neutral “how do I…?” queries demonstrating whether ChatGPT actually substitutes wikiHow wording or structure when the source is not named. That could become particularly important because the existing OpenAI multidistrict litigation is already examining large samples of ordinary user prompts and outputs. As of August 2026, the court is still dealing with summary-judgment and expert-evidence scheduling rather than having definitively resolved the overarching fair-use question.
There is another legal complication peculiar to wikiHow. Copyright does not protect an idea, procedure, process, system or method of operation. 17 U.S.C. §102(b) expressly says so. Nobody can own the idea that a shipping label should be placed in a particular location or the underlying procedure for restringing a guitar. wikiHow must therefore show copying of protectable expression: wording, explanatory choices, illustrations, original combinations and sufficiently creative selection or arrangement.
That makes verbatim reproduction potentially powerful evidence, but generic summaries of purely functional instructions considerably less so.
The fair-use argument cannot currently be treated as settled in either direction. Section 107 requires courts to weigh purpose and character, the nature of the work, the amount used and market effect.
The California AI cases point in different directions even where defendants ultimately prevailed. In Bartz v. Anthropic, Judge Alsup held that using lawfully obtained books to train Claude was highly transformative and fair use, while separately holding that Anthropic could not justify its pirated central library merely by pointing to a later training purpose. In Kadrey v. Meta, Meta also obtained summary judgment, but Judge Chhabria stressed that the plaintiffs had failed to supply sufficient evidence of market harm and warned that generative AI could, in other cases, severely undermine markets for the works on which it was trained.
wikiHow potentially has a better fourth-factor story than those plaintiffs. A novel and a chatbot answer do not necessarily fulfil the same consumer demand. But a wikiHow article entitled “How to Put a Shipping Label on a Box” and a ChatGPT answer telling someone how to put a shipping label on a box quite plainly can.
Conversely, the functional and factual nature of wikiHow’s material gives OpenAI a stronger argument under fair-use factor two than it would have with highly expressive fiction.
The most important difference is that training is only one layer of this case.
The court could conceivably conclude that statistical model training is transformative while reaching a different conclusion about RAG. Fetching a current article, inserting it into a prompt and using it to answer the exact information request for which that article was created looks much closer to traditional content reuse.
Second, wikiHow has an unusually direct substitution theory. ChatGPT is not merely producing another work in the same genre. It may answer the exact transactional question that would have generated wikiHow’s page view. That makes the case conceptually closer to parts of New York Times v. OpenAI than to the book-training cases, but arguably with an even shorter substitution chain.
Third, wikiHow’s CMI allegation is unusually specific. This matters because the Southern District of New York has already distinguished weak CMI pleadings from stronger ones. In the Times litigation, the court dismissed the Times’s §1202(b)(1) theory against OpenAI where the absence of CMI in excerpts did not adequately demonstrate that OpenAI removed CMI from its training copies, while allowing more factually detailed CMI claims by other news plaintiffs to proceed. wikiHow has clearly drafted around that problem by naming extraction technologies, describing how they allegedly discard titles, bylines and copyright information, and connecting them to particular datasets. A later decision involving Ziff Davis similarly allowed a detailed §1202(b)(1) claim against OpenAI to proceed.
Fourth, the timing is important. wikiHow registered the asserted works in August 2025 and has deliberately pleaded substantial post-registration crawling, retrieval and output activity. That appears designed to counter restrictions on statutory damages for infringement commencing before registration.
Finally, the output claim arrives in a relatively favourable procedural environment. In October 2025, Judge Sidney Stein in the OpenAI copyright MDL rejected OpenAI’s attempt to dismiss authors’ output-based infringement claims, finding that at least some ChatGPT outputs could reasonably be regarded as substantially similar to protected works. Importantly, however, that ruling expressly did not decide fair use.
ChatGPT’s expectation is therefore not an all-or-nothing judgment that “AI training is legal” or “AI training is copyright infringement.” The more plausible outcome is legal separation of the different stages of the AI pipeline.
The core claims are likely strong enough to get substantially beyond the pleading stage. The output claim in particular benefits from the existing SDNY precedent. The detailed CMI allegations also have a reasonable prospect of surviving an early challenge.
At summary judgment, however, the case could fragment. A court might decide that using lawfully accessible web material for genuine model training was sufficiently transformative to qualify as fair use while simultaneously finding that particular verbatim outputs were infringing. It could also distinguish training from ongoing RAG and conclude that unlicensed retrieval and presentation of publisher content requires a different fair-use analysis.
The RAG issue may ultimately be the most commercially important part of the case. It raises a question that will become increasingly central as AI shifts from static foundation models toward agents, search systems and continuously grounded services: even if a developer may lawfully learn from a work, does that entitle it continually to retrieve the work in order to operate a competing information service?
A negotiated settlement is also very plausible. Such a settlement could combine compensation, an AI-content licence, crawler restrictions, attribution requirements, limits on verbatim reproduction and agreed treatment of wikiHow material in retrieval systems. That would allow both parties to avoid the risk of an unpredictable precedent.
The complaint’s requested injunction is much more ambitious. wikiHow asks the court to prevent OpenAI from operating models trained on wikiHow works without authorisation. A court ordering an entire major model to cease operating because some training data included wikiHow would be extraordinary. A narrower remedy—cessation of further scraping, exclusion from RAG, removal of particular repositories, output safeguards or licensing—is considerably more plausible.
Damages could nevertheless be significant if willful post-registration infringement is eventually proven. But spectacular calculations based simply on multiplying 11,211 articles by the maximum $150,000 statutory figure should be treated cautiously. Questions about registration timing, what constitutes a separate “work” for statutory-damages purposes, fair use, willfulness and which individual acts are actionable would all have to be resolved.
The wider lesson from this lawsuit is that an AI developer should stop treating copyright compliance as a single question—“Can we train on this?”—and instead create rights controls across the entire lifecycle of content.
Build provenance into the data pipeline. Every ingested document should retain source URL, publisher, author, retrieval date, licence status, applicable restrictions, copyright notices and the purpose for which the copy may be used. Provenance should survive chunking, embedding and transformation rather than being discarded during extraction.
Separate permissions by use case. Training, model evaluation, search indexing, RAG, summarisation, output display and verbatim quotation are different activities. A licence or legal basis for one should not automatically be interpreted as permission for all the others.
Respect rights reservations technically. Training crawlers, search crawlers, user-initiated browsing agents and RAG crawlers should have distinct identities and purposes. robots.txt and publisher restrictions should be machine-enforced against those purposes. Where a commercial deal covers only specified content, access should be limited to agreed domains, paths, objects or content IDs rather than giving a general-purpose bot unrestricted access to the publisher.
Preserve CMI instead of deliberately stripping it. Cleaning HTML does not require destroying provenance. Body text can be separated from navigation while title, author, publisher, copyright, source URL and rights information are retained as structured metadata. Doing so dramatically reduces the kind of §1202 argument wikiHow is now making.
Design RAG around authorisation. A publicly accessible webpage should not automatically become an unrestricted grounding source. Retrieval indexes should know whether a source permits discovery only, short quotation, summarisation, grounding, commercial answers or model training.
Test for economic substitution, not merely memorisation. Conventional evaluations ask whether a model can reproduce copyrighted text. Developers should also ask: Can the AI satisfy the user’s entire informational need using this publisher’s material so that the user no longer needs the publisher? That is particularly important for reference works, instructional content, news, dictionaries, professional information and scholarly material.
Introduce source-sensitive output controls. A request such as “give me the verbatim text from [named copyrighted article]” should trigger quotation limits, citation, linking or refusal unless the source licence authorises reproduction. Similarity detection should operate against high-value licensed and protected corpora.
Maintain auditable content lineage and deletion capability. A developer should be able to answer a rights holder’s basic questions: Was this work collected? When? By which crawler? For what purpose? Which dataset contains it? Is it in a retrieval index? Which models used that dataset? Can further use be stopped? Systems that cannot answer those questions turn relatively manageable rights disputes into litigation.
Respond differently once a rights holder objects. Continuing large-scale collection after explicit notice, technical exclusion and licensing discussions creates a very different factual record from an inadvertent historical crawl. A sensible governance process should automatically escalate disputed content, suspend questionable new ingestion and create a documented legal/business decision before collection resumes.
Do not use the model itself as a provenance oracle. An LLM saying “I was trained on X” or “I wasn’t trained on X” should never be treated as evidence. Training provenance must come from auditable records outside the model.
The most significant aspect of wikiHow v. OpenAI is therefore not that another publisher has sued an AI company over training data. It is that wikiHow has attempted to connect the entire technical and commercial chain: identifiable crawling, notice and robots restrictions, training, runtime retrieval, reproducible output, discarded attribution, licensing opportunities and loss of the very web traffic that finances the underlying works.
Some parts of that chain are substantially better supported than others. The complaint currently overreaches when it treats crawler activity as proof of training, relies on ChatGPT’s own statements about its training history, assumes new models necessarily re-ingested wikiHow, or extrapolates from specially engineered verbatim prompts to ordinary user behaviour. Those gaps will require discovery and empirical evidence.
Nevertheless, the case may pose a harder problem for OpenAI than a pure training-data lawsuit. The strongest wikiHow theory may ultimately not be “our articles helped teach the model to write.” It may be “your service repeatedly retrieves our articles, removes their provenance, gives users substantially the answer they came to us for, and uses our own investment to replace the visit that pays for producing it.”
If the evidence ultimately supports that proposition, the court will have to confront something broader than the legality of AI training. It will have to decide where lawful machine learning ends and an unlicensed, AI-mediated substitute for a publisher begins.
No posts

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.