NVIDIA Blackwell GPUs achieve 98% performance while running encrypted AI
NVIDIA has released technical benchmarks demonstrating that its Confidential Computing technology, when enabled on Blackwell GPUs, maintains up to 98% of native inference performance. The data aims to show that high-speed, hardware-encrypted AI processing can meet regulatory privacy requirements for production-scale streaming and inference workflows with minimal latency overhead.
Key Takeaways
- Tests on the HGX B300 with Qwen 3.5-397B-A17B-FP8 maintained up to 98% of standard throughput.
- Hardware-level security features include fused private signing keys and NVIDIA Remote Attestation Service (NRAS) verification.
- Performance optimizations in SGLang and FlashInfer mitigate latency by using async copy workers and piecewise CUDA graphs.
- Confidential Computing overhead for median time per output token (TPOT) remained typically under 8% across varying batch sizes.
Why It Matters
This performance parity removes the primary barrier to adopting hardware-encrypted AI: the 'security tax' on speed. For streaming and media companies handling sensitive user data or proprietary IP, hardware-rooted security allows for production-scale inference without the latency trade-offs that previously hindered real-time applications. As regulatory pressure regarding data sovereignty increases, this capability ensures high-throughput AI workflows remain compliant. Watch for whether cloud service providers prioritize Blackwell-based confidential VMs in their upcoming 2026 infrastructure refreshes to capture enterprise workloads.
Additional Context
The push for secure AI execution environments follows a broader industry shift toward confidential computing as a standard for high-stakes enterprise workloads. Per Fortune, March 2026, the global demand for AI-specific security hardware has risen as financial and healthcare sectors move past the pilot phase into full-scale deployment of generative models. NVIDIA’s focus on the Blackwell architecture aligns with competitive moves from AMD, which recently updated its EPYC processors to enhance SEV-SNP (Secure Encrypted Virtualization) capabilities. Market analysts from Gartner noted in May 2026 that the ability to protect model weights from infrastructure providers is becoming a non-negotiable requirement for high-value B2B AI services. Recent shifts in the regulatory landscape also frame these technical milestones. According to a June 2026 report from The Wall Street Journal, the European Union's latest AI safety guidelines emphasize hardware-level data isolation as a preferred method for meeting strict sovereignty standards. Furthermore, the integration of tools like SGLang and FlashInfer into the NVIDIA stack highlights a maturing software ecosystem designed to handle the complexities of mixed-batch inference. As reported by TechCrunch in April 2026, the rise of 'agentic AI'—where autonomous systems handle sensitive customer interactions—requires the exact combination of low latency and verifiable integrity that the Blackwell benchmark aims to prove.
Read full article at developer.nvidia.com
Get this in your inbox → Subscribe
Enjoy our coverage?
Add StreamingMeme as a preferred source on Google to see more of our streaming news at the top of your Search results.
Add as preferred source