Google develops 'Frozen v2' chip to hardwire Gemini architecture into silicon
Google is reportedly developing a custom AI chip codenamed 'Frozen v2' optimized for its Gemini model architecture, with a planned rollout to data centers in 2028. The processor aims to achieve significant efficiency gains by minimizing data movement and optimizing on-chip memory usage for AI workloads.
Key Takeaways
- Frozen v2 architecture aims to reduce calculations by using operator fusion and hardwiring specific Gemini model circuits.
- The chip targets a 6x to 10x efficiency gain in tokens served per unit of power compared to current Google silicon.
- On-chip memory optimization intends to solve data bottlenecks by keeping neural network elements directly on the processor.
- Google plans a parallel family of processors, with Frozen v2 complementing rather than replacing existing Tensor Processing Units (TPUs).
- Data center deployment is scheduled for 2028, pending final engineering decisions on model-logic hardcoding.
Why It Matters
The shift from general-purpose AI accelerators to model-specific silicon marks a critical phase in infrastructure consolidation. By 'freezing' Gemini's architecture into hardware, Google can bypass the energy efficiency ceiling of traditional GPUs, which must support multiple model types. This vertical integration provides a major cost advantage for high-volume inference, potentially lowering the barrier for real-time video processing and multimodal streaming applications. For the broader ecosystem, it signals a move away from the merchant silicon bottleneck as hyperscalers trade architectural flexibility for massive operational savings. Watch for Google to prioritize those Gemini-native workloads that currently strain cloud capacity, such as 24/7 metadata generation and dynamic ad insertion.
Additional Context
The development of Frozen v2 arrives as Google aggressively expands its custom silicon portfolio to counter Nvidia's dominance. In April 2026, Broadcom confirmed it reached an expanded long-term agreement with Google to supply custom AI chip components through 2031, per CNBC. This followed the general availability of Ironwood (TPU v7) in late 2025, which Google industry analysts at Intuition Labs reported delivering a 4x performance jump over previous Trillium chips. These hardware cycles are increasingly vital as Google faces internal compute shortages that have reportedly forced its cloud division to reject certain enterprise deals, according to The Information in July 2026. Google is also diversifying beyond pure AI accelerators with general-purpose custom hardware. Per Shacknews, Google's Arm-based Axion CPU reached general availability in late 2024 to power high-traffic internal services like YouTube Advertising and BigQuery. This multi-pronged hardware strategy mirrors shifts at rival hyperscalers; Microsoft announced its second-generation Maia 200 AI chip in January 2026, specifically targeting inference efficiency for OpenAI's GPT models, according to GeekWire. Meanwhile, Amazon is exploring merchant sales of its Trainium processors as of June 2026 to capitalize on external data center demand, per Bloomberg. Technically, the 'Frozen' approach signals a move toward Application-Specific Integrated Circuits (ASICs) that are even more specialized than standard TPUs. Whereas TPUs are flexible across various neural network architectures, Frozen v2 seeks to optimize Gemini’s specific math operations. This 'software-to-silicon' pipeline is supported by a deepenening relationship with Samsung; according to a July 2026 report from CryptoBriefing, Google is in talks with Samsung to use its 2nm process for a related component codenamed Icefish, aimed at further reducing power consumption across its next-generation AI server racks.
Read full article at siliconangle.com
Get this in your inbox → Subscribe
Enjoy our coverage?
Add StreamingMeme as a preferred source on Google to see more of our streaming news at the top of your Search results.
Add as preferred source