EDPB issues draft GDPR rules for generative AI data scraping
The European Data Protection Board has released draft guidelines detailing GDPR compliance requirements for organizations scraping data to train generative AI models. The document clarifies the application of legal bases, data minimization, and transparency obligations, specifically addressing the challenges of large-scale automated data collection.
Key Takeaways
- Legal basis for large-scale scraping must rely on 'Legitimate Interest' rather than consent, which the EDPB deems unworkable in practice.
- Data minimization principles require developers use filters to exclude sensitive data, high-risk sites, and minors' information before collection begins.
- Special category personal data (SPD) collection is not automatically unlawful if it is incidental and the controller has implemented robust technical safeguards.
- AI developers using pre-scraped third-party datasets bear independent controller responsibility and must conduct upstream due diligence on data provenance.
Why It Matters
The guidelines establish a stricter compliance floor for the automated data collection that powers large language models (LLMs). By treating robots.txt and ai.txt signals as relevant indicators of user expectations, the EDPB is effectively codifying technical opt-outs into the GDPR balancing test. This shift forces video and digital media platforms to audit their scraping pipelines or face potential enforcement as 'joint controllers.' For the broader ecosystem, it signals that the era of treating the open web as a free, unregulated training corpus is ending. Watch for how major AI labs adjust their data acquisition strategies before the public consultation period closes on October 30, 2026.
Additional Context
The EDPB’s draft guidelines arrive amid a period of intense regulatory and legal pressure on AI developers regarding data provenance. Per Reuters, in March 2026, an Italian court overturned a €15 million fine previously issued by the country's data protection authority against OpenAI, though the ruling was based on procedural grounds rather than a total clearance of OpenAI's data practices. This follows high-profile enforcement actions elsewhere, including a $33.7 million fine against Clearview AI by the Dutch Data Protection Authority in September 2024 for illegal biometric data harvesting via web scraping.
Simultaneously, the technical landscape for opting out of AI training has hardened. In May 2025, per PrivacyGuides, Meta was forced to notify EU users of their right to opt out of AI training on public posts, leading to a temporary pause in Meta’s AI rollout across the region following complaints from the advocacy group NOYB. This regulatory scrutiny is now converging with the EU AI Act, which enters full enforcement in August 2026. Per DataSostech, July 2026, the AI Act adds a secondary compliance layer requiring general-purpose AI model providers to publish summaries of their training data and respect machine-readable opt-outs like the TDM (Text and Data Mining) reservation.
Media organizations are increasingly leveraging these evolving legal frameworks to protect their proprietary content. According to Tendem.ai, May 2026 reports indicated that nearly 70% of generative AI models remain reliant on scraped web data, prompting a wave of class-action lawsuits from content creators against firms including Nvidia and Snap. The EDPB guidelines provide a long-awaited bridge between these copyright disputes and fundamental data protection rights, clarifying that AI developers cannot bypass GDPR obligations simply by sourcing data through third-party brokers or automated crawlers.
Read full article at google.com
Enjoy our coverage?
Add StreamingMeme as a preferred source on Google to see more of our streaming news at the top of your Search results.
Add as preferred source