Snowflake and Google Cloud Link via Apache Iceberg for Data Federation
Snowflake and Google Cloud have announced a partnership to enable bidirectional data federation between their respective platforms using Apache Iceberg and the Iceberg REST Catalog specification. This integration allows users to access shared data across engines without ETL, and supports AI agents through the Model Context Protocol.
Key Takeaways
- Unified governance is enforced across catalog boundaries using Snowflake Horizon and Google Cloud’s Lakehouse runtime catalog.
- Support for the Model Context Protocol (MCP) enables AI agents, including Gemini Enterprise and Snowflake CoWork, to programmatically access shared data.
- Snowflake Horizon now embeds standards-compliant IRC endpoints to manage RBAC, masking, and data lineage for external engines like BigQuery.
- The integration uses the Apache Polaris open-source catalog to facilitate multi-engine interoperability without proprietary adapters.
Why It Matters
This integration effectively neutralizes the primary friction point in multi-cloud streaming architectures: data gravity. By standardizing on the Iceberg REST Catalog, media companies can now run specialized AI workloads in Google Cloud while maintaining enterprise governance in Snowflake, all without the billable overhead of moving petabytes of content metadata. In the broader ecosystem, this signals a shift from platform lock-in to 'catalog wars,' where the value resides in the governance layer rather than the storage format. Watch for broader adoption of the Model Context Protocol (MCP) by third-party AI agents seeking standardized access to these federated lakehouses.
Additional Context
The collaboration marks a critical consolidation around Apache Iceberg, which has transitioned from a Netflix engineering project to a de facto industry standard for open lakehouse architectures. Per InfoQ, Google Cloud officially rebranded its BigLake suite to 'Lakehouse for Apache Iceberg' in April 2026, signaling a full commitment to the open format for its analytics and AI pipelines. This alignment follows a period of intense competition between major data players; for instance, Databricks acquired Tabular in late 2024 to bolster its own Iceberg capabilities. A key driver for this specific integration is the maturation of the Model Context Protocol (MCP). According to Anthropic, which open-sourced the protocol in November 2024, MCP acts as the 'USB-C' for AI, allowing agentic workflows to discover and query enterprise data through one standard interface rather than hundreds of custom connectors. By mid-2025, major providers including OpenAI and Microsoft had adopted the standard, setting the stage for the deep AI-to-data integration currently being deployed by Snowflake and Google. Furthermore, the release of Apache Polaris as a top-level project by the Apache Software Foundation in early 2026 has provided the vendor-neutral foundation necessary for this level of interoperability. Snowflake and Dremio originally co-created Polaris to solve the challenge of cross-catalog metadata management. Its widespread adoption by Google Cloud and AWS (via Glue) has pressured the market to provide managed IRC endpoints, reducing the operational burden on data teams that previously had to maintain DIY metadata stores to achieve cross-engine consistency.
Read full article at snowflake.com
Enjoy our coverage?
Add StreamingMeme as a preferred source on Google to see more of our streaming news at the top of your Search results.
Add as preferred source