Nvidia Groq 3 LPX production begins to accelerate agentic AI workloads
Nvidia has launched its Groq 3 LPX inference accelerator into full production, utilizing technology licensed from Groq Inc. to enhance agentic AI workloads. The accelerator integrates with Nvidia's Vera Rubin platform and has been adopted by cloud provider Nebius Group to improve token generation speeds.
Key Takeaways
- Groq 3 LPX achieved 3,400 tokens per second running the Gemma 4 31B model in benchmark tests.
- Nebius Group is the first cloud provider to deploy the new chips via its Token Factory platform.
- A single rack-scale deployment can support up to 256 LP30 accelerators linked by high-bandwidth interconnects.
- Nvidia paid $20 billion to license the underlying technology and hired Groq founder Jonathan Ross.
- SpaceX plans to utilize the Vera Rubin platform for CPU-intensive orchestration across terrestrial and orbital systems.
Why It Matters
The move into full production signals a shift from general-purpose GPU compute toward specialized inference architectures optimized for autonomous agents. By offloading token generation from the primary Vera Rubin GPUs, Nvidia is addressing the 'decode latency' bottleneck that currently hinders real-time reasoning in complex AI workflows. For the streaming and broader tech ecosystem, this suggests a future where AI-driven personalization and metadata tagging move from batch processing to near-instantaneous execution. As cloud providers like Nebius Group integrate these accelerators, the industry should watch for a decrease in the cost-per-token for high-context agentic AI content services, which could lower the barrier for deploying sophisticated AI assistants at scale.
Additional Context
Nvidia's decision to bring Groq 3 LPX into full production places it in direct competition with a rapidly expanding field of inference-focused silicon. Groq Inc. had already begun shipping its own LPU-based inference cloud to enterprise customers by early 2026, offering per-token pricing that undercut major hyperscalers on large-context workloads. The licensing arrangement between Nvidia and Groq Inc. represents an unusual collaboration between what had been competing approaches to inference acceleration, and Jonathan Ross, Groq's CEO, described the deal as validation that deterministic execution architectures outperform general-purpose GPUs for decode-heavy workloads. Nebius Group, the cloud provider spun out of Yandex and led by CEO Danila Shtan, is among the first to integrate the accelerator into its managed inference platform, positioning it against AWS Inferentia and Google's TPU-based offerings.
The business implications extend beyond raw throughput. Nebius Group reported in its Q2 2026 earnings that inference workloads now account for 62% of total GPU cloud revenue, up from 41% a year earlier, reflecting a broader industry shift from training-dominated spending toward inference at scale. Nvidia's Vera Rubin platform, which pairs the Groq 3 LPX with next-generation GPU dies, received its first hyperscaler commitment when SpaceX's Starlink division announced plans to use the architecture for onboard satellite edge inference, a deployment that would push inference hardware into non-datacenter environments. The Gemma 4 31B model from Google DeepMind was cited as a reference workload during the launch, suggesting that open-weight models are becoming primary targets for specialized inference hardware rather than proprietary frontier models alone.
On the technical side, the Groq 3 LPX's architecture disaggregates prefill and decode phases across separate silicon, a design pattern that multiple vendors are now pursuing. Cerebras Systems published benchmark data in July 2026 showing its WSE-3 achieving 2,100 tokens per second on Llama 3.1 70B, a figure that Groq Inc.'s own LPU cloud had previously matched at smaller model sizes. Nvidia's entry into this space with a dedicated decode accelerator signals that the company sees inference disaggregation as essential to its next-generation platform strategy rather than an optional optimization. For streaming and video applications specifically, the reduction in per-token latency could enable real-time content understanding pipelines that previously required batch processing, though analysts at SemiAnalysis noted that the Groq 3 LPX's advantage narrows significantly for workloads with context windows below 32,000 tokens, which covers many current video metadata and recommendation use cases.
Read full article at siliconangle.com
Enjoy our coverage?
Add StreamingMeme as a preferred source on Google to see more of our streaming news at the top of your Search results.
Add as preferred source