Google develops Frozen v2 chip to embed Gemini architecture into silicon
Alphabet is developing a custom server chip, codenamed Frozen v2, to improve the power efficiency of its Gemini AI models. The project, expected for 2028, forms part of a broader strategy to optimize infrastructure and reduce reliance on third-party hardware as AI-related capital expenditures increase.
Key Takeaways
- Frozen v2 estimates efficiency at 6-10 times current tokens-per-watt following internal data-center capacity crunches.
- Hardware design 'freezes' specific Gemini neural-network logic into silicon while maintaining updatable model weights.
- Alphabet's 2026 capital expenditure guidance remains at $180 billion to $190 billion to fund infrastructure buildouts.
- Projected 2028 launch creates a new specialized silicon line, complementing rather than replacing existing Tensor Processing Units.
Why It Matters
The shift from general-purpose accelerators to model-specific silicon signals a vertical integration pivot to stabilize the unit economics of AI inference. By etching model logic into metal, Google aims to bypass the power-intensive data movement inherent in existing architectures, a move necessary to justify its nearly $190 billion annual infrastructure spend. This strategy forces a trade-off between extreme efficiency and future flexibility, potentially locking Google into current model architectures for years. Watch for whether external cloud customers can access these specialized chips or if they remain exclusive to internal Gemini workloads to resolve internal supply bottlenecks.
Additional Context
The development of Frozen v2 arrives as high-scale AI labs move to mitigate the rising costs of LLM inference. Per OpenAI in June 2026, the company partnered with Broadcom to tape out its own inference processor, Jalapeño, in just nine months. Unlike Google’s 2028 timeline, OpenAI expects Jalapeño to arrive in data centers by late 2026. Simultaneously, Anthropic began exploratory talks with Samsung in July 2026 to manufacture custom 2-nanometer chips, following its massive Series H funding round where Samsung participated as an infrastructure partner. This shared industry roadmap highlights a collective intent to reduce total reliance on Nvidia’s general-purpose H-series and B-series chips. Google’s own internal pressure has reportedly reached a critical threshold. Per The Information and subsequent reports in July 2026, Google Cloud has been forced to reject some external enterprise orders due to a internal capacity crunch, even reportedly agreeing to a $1 billion monthly compute lease with SpaceX to bridge the gap. While Google previously launched the ARM-based Axion CPU in April 2025 to handle general-purpose cloud tasks like YouTube ads and BigQuery, Frozen v2 represents a deeper technical commitment to optimizing the 'full stack' specifically for generative AI workloads. This hardware-level specialization is intended to protect operating margins as model sizes grow and electricity becomes the primary cost bottleneck.
Read full article at techcrunch.com
Get this in your inbox → Subscribe
Enjoy our coverage?
Add StreamingMeme as a preferred source on Google to see more of our streaming news at the top of your Search results.
Add as preferred source