Major publishers sue Google over Gemini AI training and copyright removal
A group of publishers and authors has filed a class action lawsuit against Google, alleging the company unauthorizedly used copyrighted materials to train its Gemini AI models. The plaintiffs claim Google bypassed search-intended permissions to access content and intentionally removed copyright metadata during the training process.
Key Takeaways
- Lawsuit alleges Google misappropriated works from Google Books and Google Play for AI training beyond agreed search-only purposes.
- Plaintiffs include Hachette Book Group, Cengage Learning, Elsevier, and author Scott Turow.
- Internal Google documents cited in the filing reportedly projected potential copyright fines between $10 billion and $100 billion.
- The suit claims Google used "known pirate sources" and removed copyright identifiers in violation of the Digital Millennium Copyright Act.
Why It Matters
This case shifts the AI copyright battleground to the Southern District of New York, distancing the legal debate from recent pro-tech 'fair use' rulings in California. For the streaming and media ecosystem, the focus on metadata removal and breach of existing data partnerships (like Google Books) threatens the 'transformation' defense typically used by AI developers. If the court sides with publishers, it could force a massive valuation reset for AI models built on scraped data, mandating localized licensing deals that would significantly increase operational costs for Big Tech.
Additional Context
The litigation against Google follows an intensifying wave of legal challenges to AI training practices across the media landscape. per AP News and Reuters (July 2026), The New York Times and a coalition of nearly 400 newspaper publishers recently escalated their separate suit against OpenAI, seeking court sanctions. These publishers allege that OpenAI misled the court regarding its technical ability to search training datasets and intentionally deleted billions of relevant ChatGPT conversation logs to obstruct discovery. The Times has reported spending over $28 million on its AI litigation through mid-2026, highlighting the high financial stakes of these intellectual property disputes. While previous rulings in California, such as those involving Meta and Anthropic in June 2025, initially favored AI companies, a significant carve-out has emerged concerning pirated content. In the Bartz v. Anthropic case (settled for $1.5 billion in late 2025), a federal judge ruled that while training on legally acquired books was transformative, maintaining a permanent research library sourced from pirate sites like Library Genesis was inherently infringing. Per The Bookseller and Publishing Perspectives (July 2026), Hachette and Cengage have also filed a parallel suit against Meta over its Llama models, indicating an industry-wide strategy to consolidate claims in New York courts where the specific nuances of 'market dilution' and 'substitutive outputs' may receive stricter scrutiny than in earlier Silicon Valley-based venues.
Read full article at techcrunch.com
Get this in your inbox → Subscribe
Enjoy our coverage?
Add StreamingMeme as a preferred source on Google to see more of our streaming news at the top of your Search results.
Add as preferred source