IETF AI scraping standards debate intensifies as EU mandates loom
The IETF AI Preferences Working Group is meeting in London to develop a technical standard for website publishers to signal scraping preferences to AI bots. The Center for Democracy & Technology is participating to ensure the standard remains voluntary and does not inadvertently block beneficial public interest data collection as it faces pressure from EU policy mandates.
Key Takeaways
- The AI Preferences Working Group aims to create a machine-readable taxonomy for specific AI data uses beyond simple binary blocking.
- EU AI Office plans to codify the IETF standard into its Code of Practice for General Purpose AI Models, effectively creating a legal compliance requirement.
- New York's Stealth Crawler Prohibition Act and a federal companion bill in the U.S. House threaten to de-anonymize automated data collection.
- Intermediaries like Cloudflare are deploying proprietary bot-blocking tools, creating a fragmented 'pay-per-scrape' environment for researchers and journalists.
Why It Matters
The transition from voluntary technical signals to legally enforceable mandates threatens the 'unspoken bargain' of the open web. If these standards become rigid, essential public interest functions—including investigative journalism, cybersecurity threat detection, and academic research—could be blocked by publishers seeking to monetize all automated traffic. This shift risks creating a 'Gated Web' where high-quality data is restricted to entities capable of paying for access, while the public web is left with low-quality AI-generated content. Industry observers should monitor the IETF London meeting outcomes to see if consensus language can preserve exceptions for non-commercial, pro-social scraping before the EU finalize its Code of Practice of Practice.
Additional Context
Cloudflare has been building commercial tooling that operationalizes the kind of bot-preference signals the IETF AI Preferences Working Group is standardizing. In 2026, Cloudflare launched its Attribution Business Insights dashboard for Bot Management customers, which provides publishers with crawl-to-referral ratios per bot operator, bandwidth consumption data, and a taxonomy that classifies crawlers as Training, Search, or Agent based on their behavior. The product is designed to give website owners quantitative leverage in licensing negotiations with AI companies, directly complementing the voluntary signaling framework under development at the IETF.
The regulatory backdrop adds urgency to the technical work. The EU's AI Act Code of Practice, which addresses transparency obligations for general-purpose AI models, is expected to reference how training data is sourced and whether publishers' machine-readable preferences were respected. The Center for Democracy & Technology has mapped the debate over AI preferences, noting that the IETF standard's voluntary nature could be undermined if EU policymakers treat compliance with bot signals as a de facto legal requirement rather than a best practice. Public Knowledge and Internet Archive have similarly argued that overly broad blocking signals could harm archival and research functions that depend on automated access.
The traffic dynamics driving this debate are measurable. Ericsson's June 2026 Mobility Report found that 43 out of 55 service providers experienced higher uplink growth than downlink, with AI-driven applications and user-generated content cited as primary drivers. Ericsson's scenario modeling projects uplink traffic could be three times higher in 2031 compared with 2025 as agentic AI workloads generate continuous data streams. While these figures describe mobile network traffic rather than web crawling specifically, they illustrate the broader infrastructure pressure from AI data flows that makes publisher-side signaling standards increasingly consequential for content economics.
Read full article at cdt.org
Enjoy our coverage?
Add StreamingMeme as a preferred source on Google to see more of our streaming news at the top of your Search results.
Add as preferred source