Etched reaches $10.3 billion valuation to scale purpose-built AI inference hardware
Inference hardware startup Etched has raised $300M in a Series C funding round led by Sequoia Capital to scale production of its architecture-agnostic AI compute clusters. The company is developing its own rack-scale systems designed to improve power efficiency and latency for large language models and non-transformer architectures.
Key Takeaways
- Series C funding led by Sequoia Capital includes participation from a16z, Jane Street, and SK Hynix.
- Company is scaling production via a Taiwan factory and a new 80,000-square-foot facility in Milpitas.
- Hardware architecture utilizes Low Voltage Inference (LVI) and Cluster Scale Memory (CSM) to bypass thermal and memory bottlenecks.
- Systems are designed to be architecture-agnostic, supporting MoE models like DeepSeek alongside state space models.
- Etched now employs 400 people, recruiting from Nvidia, Broadcom, Google TPU, and SK Hynix.
Why It Matters
The massive valuation jump — doubling to $10.3 billion in seven months — signals investor conviction that the AI infrastructure bottleneck is shifting from training to inference. By building specialized ASICs rather than general-purpose GPUs, Etched aims to provide a high-throughput, lower-cost alternative for hyperscalers and enterprises struggling with the power demands of frontier models. In the streaming and B2B video ecosystem, this could lower the barrier for real-time generative video and agentic search features that currently face prohibitive compute costs. Watch for the delivery of the first rack-scale units this summer to verify if the hardware meets its efficiency claims in production environments.
Additional Context
Etched’s rise comes as the industry increasingly seeks specialized alternatives to Nvidia’s general-purpose H100 and Blackwell architectures. Per TechPowerUp and Tom’s Hardware (June 2024), Etched’s flagship Sohu chip is an ASIC hard-coded for transformer attention. The company claims a single eight-chip Sohu server can process 500,000 tokens per second on Llama-3 70B, which would represent a significant throughput advantage over standard GPU clusters. This performance leap is attributed to using nearly 90% of the silicon for matrix multiplication, compared to approximately 3.3% in traditional GPUs.
The startup’s commercial momentum is backed by a reported $1 billion in signed customer contracts, according to EE Times (July 2026). While Etched has not publicly named its early customers, its vertical integration strategy — encompassing a 2 MW data center in San Jose and a 10 MW prototyping lab in Milpitas — reflects the aggressive operational scaling required to compete with established chipmakers. TechCrunch and other outlets (July 2026) noted that this Series C is the highest valuation ever for a Sequoia-led round at this stage, following a $500 million round in late 2025.
This funding surge mirrors a broader trend in the semiconductor market where specialized silicon providers are gaining traction. For instance, per Bloomberg and VentureCapital.com (July 2026), quantitative trading firms like Jane Street and Two Sigma have participated in Etched’s rounds, suggesting that industries with extreme latency requirements are among the primary early adopters. As hyperscalers like Google and Amazon continue to iterate on their own internal TPUs and Trainium chips, Etched is positioning itself as the primary independent provider for organizations that want to optimize for the specific math of modern large language models without being locked into a single cloud provider's ecosystem.
Read full article at storagenewsletter.com
Enjoy our coverage?
Add StreamingMeme as a preferred source on Google to see more of our streaming news at the top of your Search results.
Add as preferred source