NVIDIA CUDA Toolkit 13.4 adds Windows on Arm and Rubin support
NVIDIA has released CUDA Toolkit 13.4, which introduces support for Windows on Arm and early developer access to the Rubin GPU architecture. The update also includes Multi-Process Service V3 for improved GPU resource partitioning and performance enhancements for Blackwell GPUs.
Key Takeaways
- Multi-Process Service V3 adds a scriptable CLI and cgroup-integrated memory limits for precise GPU resource isolation
- CUB DeviceScan performance on Blackwell GPUs increased from 50% to 92% memory-bandwidth utilization
- CUDA Compute Fabric Transport enables asynchronous data movement across NVLink fabric without mapping remote allocations to virtual address space
- The SDK installer no longer bundles the NVIDIA driver, requiring separate installation of nvidia-open or toolkit packages
Why It Matters
The expansion to Windows on Arm signals a significant shift in the hardware landscape for high-performance video and AI development, moving beyond traditional Linux-based Arm environments. By providing preview access to the Rubin architecture and optimizing Blackwell performance, NVIDIA is ensuring the software stack remains ahead of hardware deployment cycles for agentic AI and complex encoding tasks. This release forces a more modular approach to infrastructure management by decoupling drivers from the toolkit, reflecting a move toward containerized and orchestrated GPU environments. Watch for the general availability of Rubin support in future toolkit releases to gauge the timeline for next-generation hardware adoption.
Additional Context
NVIDIA's CUDA platform is rapidly expanding beyond its traditional Linux and x86 strongholds into new hardware and industry verticals. The company's partnership with Nokia has become a central pillar of its telco AI strategy, with Nokia's entire RAN strategy now built on its close partnership with Nvidia, cemented by the chipmaker's $1 billion investment in the Finnish company. That investment positions CUDA as the foundational software layer for GPU-accelerated AI-RAN deployments across multiple operators, including T-Mobile US, SoftBank, and Vodafone. Meanwhile, Ericsson launched its AI in RAN commercial software subscription on June 11th, claiming up to 20% higher downlink throughput and up to 10% better spectral efficiency across more than 15 live deployments using existing baseband silicon, representing a competing approach that embeds AI directly in radio hardware without requiring NVIDIA's CUDA stack.
The business implications of CUDA's expanding reach are significant for telco infrastructure economics. Nokia has announced work with AWS and Databricks to build the data, cloud, and control layers for autonomous networks, with its Autonomous Network Fabric set to run on AWS from later this year. Nokia claims its autonomous networks portfolio is already delivering automation rates higher than 90 percent, service delivery times of four hours or less, and up to 85 percent reduction in slice rollout time. This multi-cloud orchestration layer sits above the CUDA-accelerated compute tier, creating a full-stack AI automation architecture that ties telco capital expenditure directly to NVIDIA's software ecosystem. The divergence between NVIDIA-backed and proprietary approaches is sharpening competitive dynamics, with Nokia and Ericsson diverging like never before on AI-RAN strategy, as Nokia pushes RAN improvements at "software speed" while Ericsson emphasizes hardware-embedded intelligence.
Technical benchmarks from live deployments are beginning to validate the CUDA-accelerated approach for network workloads. Ericsson's AI-native scheduler produced roughly 10% better spectral efficiency and up to about 15% higher downlink throughput versus traditional rule-based schedulers in live trials with T-Mobile's 5G Advanced network, with trials expanding across Los Angeles, New York, and Salt Lake City starting in early Q2 2025. T-Mobile is targeting full commercial deployment in Q3 2026. These results demonstrate that AI-driven network optimization, whether running on CUDA-accelerated GPUs or embedded baseband processors, is moving from proof-of-concept to production scale. The CUDA Toolkit 13.4 release, with its Windows on Arm support and Rubin architecture preview, ensures that NVIDIA's developer ecosystem remains compatible with the increasingly diverse hardware environments where these AI workloads will run, from edge radios to cloud data centers.
Read full article at developer.nvidia.com
Enjoy our coverage?
Add StreamingMeme as a preferred source on Google to see more of our streaming news at the top of your Search results.
Add as preferred source