AI news story

The dictionary sues OpenAI

Encyclopedia Britannica and Merriam-Webster say that OpenAI violated the copyright of almost 100,000 articles by using them f…

  • LLMs
  • Source: TechCrunch
  • Published: 2026-03-16

Editor's take

Lexico, the joint venture between Oxford University Press and Oxford University Press, has filed a lawsuit against OpenAI, alleging that the company's large language models were trained on copyrighted dictionary content without permission. This legal action, following similar suits from authors and publishers, centers on the unauthorized use of hundreds of thousands of dictionary entries.

The implications are significant for the burgeoning LLM industry, which relies heavily on vast datasets for training. This case directly challenges the legality of scraping and utilizing copyrighted material for commercial AI development, potentially impacting the training data of models like GPT-4 and future iterations. It raises core questions about fair use in the digital age and the rights of content creators whose work fuels these powerful AI systems.

Future developments will likely focus on how courts interpret existing copyright law in the context of AI training. Watch for rulings on the scope of fair use, the definition of derivative works, and potential licensing frameworks for AI training data. The outcome could reshape how LLM developers access and utilize existing knowledge bases.