AI news story
The dictionary sues OpenAI
Encyclopedia Britannica and Merriam-Webster say that OpenAI violated the copyright of almost 100,000 articles by using them f…
Editor's take
Lexico and Merriam-Webster have filed suit against OpenAI, alleging that the company's large language models, like GPT-4, were trained on copyrighted dictionary content without permission. This legal action highlights the growing tension between AI developers' insatiable need for training data and the rights of intellectual property holders.
The implications extend beyond these two publishers. If successful, this lawsuit could set a significant precedent for how LLMs are trained, potentially forcing AI companies to secure licenses for vast swathes of published text, impacting the cost and accessibility of future models. It also raises questions about fair use in the digital age and the economic viability of content creation when that content can be freely ingested by AI.
Future developments to monitor include OpenAI's defense strategy and whether other publishers will follow suit. The outcome could either usher in a new era of licensing agreements for AI training data or lead to more restrictive interpretations of copyright law in the context of machine learning, potentially limiting the scope of what LLMs can learn.