Crusoe and Perplexity have entered a multi-year partnership to utilize Crusoe Cloud's infrastructure for Perplexity's AI model training and managed inference. The agreement leverages NVIDIA GB300 NVL72-powered clusters and InfiniBand solutions to support Perplexity's real-time research and search products.
This partnership signals a shift toward vertically integrated infrastructure where compute providers manage the entire lifecycle from power generation to token delivery. For the streaming and search ecosystem, the move to NVIDIA GB300 NVL72 clusters highlights the escalating hardware requirements needed to maintain real-time responsiveness in agentic AI applications. By moving inference to a managed service with proprietary optimizations, Perplexity aims to solve the latency bottlenecks that often plague large-scale AI deployments. As competition intensifies among AI-driven search platforms, watch for whether this infrastructure consolidation results in measurable gains in user retention or a significant reduction in per-query compute costs.
Crusoe has rapidly expanded its footprint in the AI compute market, positioning itself as a vertically integrated alternative to hyperscalers. The company's cloud platform now serves multiple AI workloads beyond Perplexity, and its partnership with NVIDIA on GB300 NVL72 clusters places it among a small group of providers deploying the latest Blackwell-generation hardware at scale. In June 2026, Ericsson launched its AI in RAN commercial software subscription claiming up to 20% higher downlink throughput across more than 15 live deployments, illustrating how AI infrastructure demand is spreading across sectors from telecom to search. Crusoe's model of owning power generation alongside GPU clusters differentiates it from pure colocation providers and aligns with the broader industry trend of AI companies seeking dedicated, low-latency compute environments rather than shared cloud capacity.
The business dynamics of managed inference are reshaping how AI companies structure their infrastructure spend. Perplexity's decision to outsource both training and inference to a single provider reflects a growing preference among AI-native firms for consolidated vendor relationships that reduce integration overhead. Nokia combined with AWS and Databricks to build a telco AI control layer at DTW Ignite 2026, demonstrating a parallel consolidation pattern in telecom where operators are pairing cloud platforms with orchestration vendors to manage autonomous network workloads. Nokia reported that operators using its autonomous networks portfolio achieved automation rates above 90 percent and service delivery times under four hours. For Perplexity, the Crusoe arrangement similarly bundles hardware, networking, and operational management into one contract, reducing the coordination burden that typically accompanies multi-vendor AI stacks.
On the technical side, the NVIDIA GB300 NVL72 architecture that underpins this partnership represents a significant step up in inference throughput per rack compared to prior-generation H100 and H200 systems. Nokia and Google Cloud unveiled six specialized AI agents at DTW Ignite 2026 capable of slashing network problem-solving times by 50% to 80%, showing how agentic AI workloads are driving demand for the same class of GPU infrastructure that Perplexity is now accessing through Crusoe. Nokia separately disclosed that its agentic AI deployment in mobile core reduced call setup times from roughly 10 seconds to one or two seconds, a latency improvement that mirrors the responsiveness gains Perplexity is targeting for its real-time search product. These benchmarks underscore why AI companies are locking in dedicated GPU capacity rather than relying on elastic cloud pricing, particularly for latency-sensitive inference workloads where even small delays affect user retention.
Perplexity has established a multi-year partnership with Crusoe Cloud to power its model training and inference using NVIDIA GB300 NVL72 clusters. This strategic move aims to reduce latency for real-time search users and consolidate infrastructure management, highlighting the industry's shift toward vertically integrated compute solutions for high-performance agentic AI applications.
Perplexity is deploying NVIDIA GB300 NVL72-powered clusters and InfiniBand solutions to support both research and production workloads.
Perplexity is adopting Crusoe’s managed inference service to leverage proprietary optimizations that improve throughput and reduce latency for its real-time search products.
As part of the agreement, Crusoe’s 1,800 employees will adopt Perplexity Enterprise Pro and Max to assist with internal data analysis.
Crusoe differentiates itself from pure colocation providers by owning power generation alongside its GPU clusters, offering a vertically integrated alternative to traditional hyperscalers.
Add StreamingMeme as a preferred source on Google to see more of our streaming news at the top of your Search results.
Add as preferred source