Photo Credit: wikiHow
DIY website wikiHow has sued OpenAI for copyright infringement over the alleged misuse of its instructional content to train the company’s popular AI chatbot ChatGPT.
On Friday, do-it-yourself site wikiHow filed a federal lawsuit in Manhattan against OpenAI with allegations that the company scraped instructional material to train its artificial intelligence models, including the popular chatbot, ChatGPT. The filing claims that OpenAI scraped thousands of its how-to articles, misusing the platform’s instructional material to train its AI to respond to human prompts.
WikiHow publishes instructional articles on anything from “routine tasks to specialized skills,” the company said in its complaint. According to the filing, OpenAI scraped more than 11,000 of its articles without permission or compensation to train its large language models (LLMs). Further, wikiHow alleges that ChatGPT reproduces its text in response to user prompts and threatens to displace its market altogether.
According to wikiHow, OpenAI copied its articles “at scale” and infringed upon at least 1,200 of the company’s registered copyrights. WikiHow is requesting unspecified damages and a court order preventing further infringement of its copyrights.
“Having ingested wikiHow’s articles, these models now produce competing how-to content on the same subjects, at a fraction of the time, effort, and cost of researching, writing, and editing a wikiHow article,” the filing reads. “The substitution cuts wikiHow’s revenue and, over time, its reason to keep producing the articles at all.”
On Monday, an OpenAI spokesperson told the press that the company’s AI models are trained on “publicly available data and grounded in fair use.”
The lawsuit is the latest brought by copyright holders, including music publishers, book publishers, artists, and authors, against AI companies over the alleged misuse of their IPs to train chatbots like OpenAI’s ChatGPT, Anthropic’s Claude, Meta’s Llama, and Google’s Gemini. The overarching theme in these lawsuits is the practice of AI tech companies hoovering up massive amounts of IP, regardless of authorization or license, before spitting it out into another format.
Many of these companies have attempted to invoke fair use in such cases. But it remains to be seen if this approach will stand as a legal defense for AI tech companies who now ask for forgiveness instead of first asking for permission.