NVIDIA has released technical optimizations for Mixture of Experts (MoE) training in JAX, utilizing its Transformer Engine to improve throughput for large-scale models like DeepSeek-V3. The update introduces grouped GEMM kernels and fused expert parallelism operations, which the company claims achieve a 10.4x performance increase on Blackwell hardware.
The 10.4x performance increase directly addresses the primary bottleneck in Mixture of Experts architectures: the communication overhead of dynamic token routing. By shifting from capacity-based constraints to a dropless MoE approach, developers can maintain higher model accuracy without the traditional compute penalties of ragged tensors. For the streaming ecosystem, these efficiencies accelerate the deployment of sophisticated recommendation engines and real-time video processing agents that rely on massive transformer models. As training costs for 600B+ parameter models remain a barrier, this throughput gain provides a viable path for specialized B2B video applications. Watch for the integration of NVFP4 quantization in future Transformer Engine releases to further drive down inference and training latency.
NVIDIA's Transformer Engine has become a central component in the company's strategy to dominate large-scale model training infrastructure. In June 2026, Ericsson launched its AI in RAN commercial software subscription claiming up to 20% higher downlink throughput across more than 15 live deployments using existing baseband silicon, illustrating how AI acceleration hardware is being deployed across telecom networks that increasingly rely on transformer-based models for optimization. The same Blackwell GPU architecture that powers NVIDIA's MoE training gains is also underpinning these network AI workloads, creating a unified hardware platform from research labs to production telecom environments. The competitive dynamics around NVIDIA's AI training stack are intensifying as rivals pursue alternative architectures. Nokia's entire RAN strategy is now built on its close partnership with Nvidia, cemented by the chipmaker's $1 billion investment in the Finnish company, with Nokia designing its Layer 1 RAN functions to run on Nvidia's CUDA software platform and GPUs. This deepening hardware dependency means that performance improvements in NVIDIA's training stack, such as the Transformer Engine's grouped GEMM kernels, directly benefit downstream partners building AI-native network functions. Nokia has combined with AWS and Databricks to build the data, cloud, and control layers for autonomous networks, positioning its Autonomous Network Fabric as an operating system for telco radio, core, transport, and service domains that will consume the kind of large-scale model training NVIDIA is optimizing. On the technical front, NVIDIA's push into agentic AI for network operations demonstrates how MoE-style architectures are being applied beyond pure training workloads. Nokia is deploying agentic AI within its mobile core, with AI-native features reducing call setup times from about 10 seconds to one or two seconds through machine learning-based paging and autonomous decision-making without human intervention. These inference-heavy workloads benefit from the same Transformer Engine optimizations that accelerate training, particularly as models grow into the hundreds of billions of parameters. Nokia went full agentic at DTW Ignite, securing NTT Docomo as a customer while Nvidia's influence looms over the vendor's roadmap, signaling that the demand for efficient large-model training and inference will continue to drive adoption of NVIDIA's acceleration stack across the streaming and telecom value chain.
NVIDIA has updated its JAX training stack with Transformer Engine optimizations, achieving a 10.4x throughput increase for Mixture of Experts models like DeepSeek-V3. By eliminating token padding and utilizing grouped GEMM kernels, the update maximizes Blackwell GPU utilization, significantly reducing training costs and bottlenecks for large-scale AI model development.
The update delivers a 10.4x throughput gain, reaching 1,068 TFLOPS per GPU on Blackwell hardware compared to the previous 103 TFLOPS baseline.
The update utilizes the Transformer Engine to eliminate token dropping and padding, while using grouped GEMM kernels to handle variable expert token counts in a single call.
Yes, the update preserves model quality by shifting from capacity-based constraints to a dropless MoE approach, which avoids the traditional compute penalties of ragged tensors.
These efficiencies accelerate the deployment of sophisticated recommendation engines and agentic AI for network operations, which rely on massive transformer models.
Add StreamingMeme as a preferred source on Google to see more of our streaming news at the top of your Search results.
Add as preferred source