Cisco Nvidia AI Factory expands to liquid-cooled rack-scale Blackwell systems
Cisco and Nvidia have expanded their Secure AI Factory partnership to include liquid-cooled, rack-scale systems utilizing Supermicro hardware and Nvidia Blackwell GPUs. The initiative aims to transition enterprise AI infrastructure from a focus on GPU acquisition to continuous, optimized token production through integrated networking and management software.
Key Takeaways
- Full rack-scale solutions featuring Blackwell GPUs and liquid cooling will be orderable starting in September
- Partnership integrates Nvidia Spectrum-X Ethernet architecture with Cisco NX-OS and SONiC operating systems
- Infrastructure metrics are shifting from GPU counts to tokens per second, tokens per watt, and tokens per dollar
- Collaboration utilizes Supermicro hardware to bridge accelerated compute with existing enterprise data environments
Why It Matters
This expansion signals a transition from the experimental GPU acquisition phase to a standardized AI production cycle. By integrating Nvidia Spectrum-X into the Cisco operating model, enterprises can now treat massive distributed clusters as a single computer rather than siloed components. This architecture is critical for streaming and media firms needing to connect proprietary data sets to inference models without bespoke supercomputing engineering. For the broader ecosystem, it establishes a validated reference design that reduces the complexity of deploying liquid-cooled hardware at scale. Watch for utilization rates and 'time to first token' to become the primary benchmarks for evaluating the ROI of these high-density rack deployments.
Additional Context
Supermicro has emerged as a critical hardware partner in the rack-scale AI factory model that Cisco and Nvidia are standardizing. In early 2026, Supermicro announced it had shipped over 100,000 liquid-cooled AI racks to hyperscalers and enterprise customers, positioning the company as one of the fastest-growing suppliers of dense GPU infrastructure. That volume underscores why Cisco chose Supermicro as the chassis layer for its Secure AI Factory reference design: the thermal and power-density requirements of Blackwell-class GPUs demand proven liquid-cooling integration at scale, and Supermicro's manufacturing throughput reduces deployment risk for enterprises unfamiliar with immersion or direct-to-chip cooling.
The competitive landscape around AI factory networking is intensifying, with Nvidia pushing its own Spectrum-X Ethernet platform as the default fabric while Cisco layers its Silicon One and NX-OS management on top. Google published new documentation in May 2026 on optimizing websites for generative AI features in Search, reflecting how AI infrastructure investments are cascading into application-layer decisions across industries. For streaming and media companies, the Cisco-Nvidia partnership matters because it provides a validated path from raw GPU capacity to production inference workloads, such as real-time content recommendation and automated metadata tagging, without requiring in-house supercomputing expertise. The rack-scale approach also aligns with broader enterprise trends toward treating AI clusters as managed services rather than capital experiments.
On the technical side, the shift toward token-based output metrics represents a measurable departure from traditional FLOPS-centric benchmarking. Deepgram's integration with AWS IAM temporary delegation provides scoped, time-bound access for support engineers directly to SageMaker endpoints, illustrating how production AI workloads increasingly demand operational governance frameworks alongside raw compute. For Cisco and Nvidia, the implication is that AI factory success will be measured not just by peak throughput but by sustained token-per-second delivery under real-world inference loads, a metric that directly maps to streaming use cases like live transcoding, content personalization, and conversational AI interfaces embedded in viewer experiences.
Read full article at siliconangle.com
Enjoy our coverage?
Add StreamingMeme as a preferred source on Google to see more of our streaming news at the top of your Search results.
Add as preferred source