Seattle Times and Newsday sue OpenAI over journalistic training data
The Seattle Times and Newsday have filed a lawsuit against OpenAI and Microsoft, alleging unauthorized use of their journalistic content to train AI models. The plaintiffs are seeking the destruction of training datasets and models that incorporate their copyrighted works, citing significant revenue losses.
Key Takeaways
- Plaintiffs are seeking the destruction of all training datasets and AI models that incorporate their copyrighted journalistic works.
- The lawsuit includes Microsoft because its Copilot assistant is built upon OpenAI technology like GPT-4.
- Nearly 400 local newspapers have joined a separate class action lawsuit against the same defendants over similar data usage concerns.
- OpenAI maintains that training models on publicly available web content constitutes fair use under U.S. copyright law.
Why It Matters
The demand for the destruction of training datasets represents a significant technical threat to the current generation of large language models. If courts reject the fair use defense, OpenAI and Microsoft could be forced to retrain models like GPT-4 from scratch at a cost of hundreds of millions of dollars. This legal pressure is already bifurcating the market between publishers who sign licensing deals, such as the Associated Press, and those pursuing litigation to protect their intellectual property. The outcome will determine whether AI companies must transition to a paid licensing ecosystem for all training inputs. Watch for whether the court grants a preliminary injunction regarding the continued use of the disputed datasets.
Additional Context
The Seattle Times and Newsday join a growing roster of publishers pursuing legal action against OpenAI over training data practices. In December 2024, The New York Times filed its own copyright suit against OpenAI and Microsoft, alleging that millions of Times articles were used without authorization to train ChatGPT and other models, a case that remains active and has become a bellwether for the industry. Meanwhile, a coalition of Alden Global Capital-owned newspapers including the Chicago Tribune and Denver Post filed suit against OpenAI and Microsoft in April 2024, expanding the geographic and ownership diversity of plaintiffs challenging AI training practices.
On the licensing side, OpenAI has pursued a parallel strategy of signing content deals to secure training rights and reduce legal exposure. OpenAI announced a licensing agreement with News Corp in May 2024 covering The Wall Street Journal, The New York Post, and other titles, reportedly valued at more than $250 million over five years. The Associated Press signed a separate licensing deal with OpenAI in July 2023, and Axios reported in November 2024 that OpenAI was in active negotiations with multiple other publishers seeking similar arrangements. This bifurcation between litigation and licensing is reshaping how news organizations approach AI companies, with some outlets treating deals as a revenue stream and others viewing them as insufficient compensation for the scale of content extraction.
The technical stakes of these lawsuits extend beyond individual damages. If courts order destruction of training datasets containing copyrighted journalism, OpenAI would face the prospect of retraining GPT-4 and successor models from filtered corpora, a process that researchers at Stanford's Center for Research on Foundation Models estimated could cost between $500 million and $1 billion in compute alone for frontier-scale models. The European Union's AI Act, which entered into force in August 2024, requires AI developers to publish detailed summaries of training data and comply with copyright opt-out mechanisms, creating a regulatory floor that could influence how U.S. courts evaluate fair use defenses in these publisher cases. The interplay between U.S. litigation outcomes and EU transparency mandates will likely determine whether a emerges for journalistic training data.
Read full article at secnews.gr
Enjoy our coverage?
Add StreamingMeme as a preferred source on Google to see more of our streaming news at the top of your Search results.
Add as preferred source