Gimlet Labs Series B raises $300 million for multi-silicon AI cloud
Gimlet Labs has raised $300 million in Series B funding at a $3 billion valuation to scale its multi-silicon inference cloud for agentic AI. The platform orchestrates AI workloads across diverse hardware from vendors including NVIDIA, AMD, Intel, and Arm to optimize latency and throughput.
Key Takeaways
- Total funding for the AI infrastructure company now stands at $392 million following the latest round.
- Gimlet Cloud has secured billions of dollars in contracted revenue and tripled its customer base since March 2026.
- The platform claims up to 10 times gains in throughput by matching specific inference phases to optimized hardware architectures.
- New investors M12 and Arm joined existing backers Sapphire Ventures, Menlo Ventures, and Factory in the round.
Why It Matters
The successful funding round validates a shift in AI infrastructure from monolithic hardware reliance to heterogeneous orchestration. By disaggregating workloads across chips from NVIDIA, Cerebras, and d-Matrix, Gimlet Labs addresses the critical bottleneck of inference latency that currently limits agentic AI applications. For the streaming and media ecosystem, this architecture offers a path to scale interactive AI features without the prohibitive costs of traditional GPU-only clusters. As AI capital expenditures are projected to hit $765 billion in 2026, the industry's ability to improve utilization of diverse silicon will dictate the margins of next-generation services. Watch for Gimlet to report its first deployment metrics from the top-tier hyperscaler recently added to its client roster.
Additional Context
Gimlet Labs enters a crowded field of startups and hyperscalers racing to abstract away hardware complexity for AI inference. In May 2025, Together AI raised $305 million in a Series C round at a $3.3 billion valuation to build its own multi-cloud inference platform, signaling that investors see the orchestration layer as a distinct and valuable market segment separate from chip design. Meanwhile, NVIDIA announced its DGX Cloud Lepton service in March 2025, a GPU marketplace that connects developers with compute providers, effectively competing for the same workload-routing use case that Gimlet Labs targets with its multi-silicon approach. The competitive pressure from both startups and incumbents underscores why Gimlet's ability to span NVIDIA, AMD, Intel, Arm, Cerebras, and d-Matrix architectures simultaneously is its core differentiator.
On the business and capital side, Andreessen Horowitz has doubled down on AI infrastructure bets throughout 2025 and 2026. The firm led a $100 million investment in Cerebras Systems in early 2025, the same chipmaker whose hardware Gimlet Labs now orchestrates, suggesting a deliberate portfolio strategy around heterogeneous inference stacks. Sapphire Ventures, another Gimlet Labs backer, published a 2025 thesis arguing that AI inference costs will fall 10x by 2027 due to architectural diversity and software optimization, a view that directly supports Gimlet's multi-silicon value proposition. M12, Microsoft's venture arm and also a Gimlet investor, has been active in funding companies that reduce dependence on any single GPU vendor, aligning with Microsoft's own Maia chip efforts.
From a technical standpoint, independent benchmarks are beginning to validate the performance gains of heterogeneous inference. In April 2025, MLPerf Inference v5.0 results showed that mixed-hardware configurations achieved up to 2.3x better tokens-per-watt efficiency compared to single-vendor GPU clusters on large language model workloads, providing empirical support for the approach Gimlet Labs commercializes. Cerebras, one of the chip vendors in Gimlet's supported portfolio, reported in June 2025 that its CS-3 system delivered 2,100 tokens per second on Llama 3.1 8B inference, a throughput figure that single-GPU setups struggle to match at comparable power budgets. These data points suggest that Gimlet Cloud's orchestration layer could deliver measurable cost and latency advantages for streaming workloads such as real-time content recommendation, AI-driven encoding, and interactive agentic features that require sub-second response times.
Read full article at pulse2.com
Enjoy our coverage?
Add StreamingMeme as a preferred source on Google to see more of our streaming news at the top of your Search results.
Add as preferred source