AWS ASR inference costs drop 75% using NVIDIA MPS optimization
AWS, NVIDIA, and Heidi Health have published a technical guide demonstrating that using NVIDIA CUDA Multi-Process Service (MPS) on Amazon EC2 G7e instances can reduce GPU infrastructure requirements for speech recognition by 75%. The architecture optimizes automatic speech recognition and diarization workloads to maintain sub-second latency while maximizing hardware utilization.
Key Takeaways
- NVIDIA CUDA Multi-Process Service allows multiple processes to share a single GPU concurrently, eliminating the 80% idle capacity typical in standard time-slicing.
- Heidi Health reduced its infrastructure requirements from 16 GPU instances to just 4 while processing 2.4 million weekly consultations.
- The optimized stack on g7e.4xlarge instances achieved 92.1 requests per second per GPU with a mean latency of 352 ms.
- Combining TensorRT and ONNX with MPS further increased throughput to 111.6 requests per second, representing an 88% infrastructure reduction.
Why It Matters
This optimization addresses the primary bottleneck in scaling AI-driven video and audio services: the high cost of underutilized GPU hardware. By partitioning a single NVIDIA L40S GPU into multiple concurrent execution contexts, streaming platforms can significantly lower the overhead of automatic speech recognition and diarization. This shift moves ASR from a high-cost infrastructure burden to a highly efficient utility, allowing smaller firms to compete with hyperscalers on transcription features. As market fragmentation continues, the ability to serve complex models with 75% less hardware will likely become the baseline for profitable AI integration. Watch for whether AWS integrates these MPS-safe CUDA graph safety mechanisms directly into its managed SageMaker inference endpoints.
Additional Context
NVIDIA's CUDA Multi-Process Service is gaining traction as a cost-reduction layer for GPU-intensive inference workloads across cloud providers. The technique of partitioning a single GPU into multiple concurrent execution contexts has become central to hyperscaler strategies for maximizing hardware utilization. In July 2026, Meta Platforms and BlackRock announced plans to build a 1-gigawatt data center complex in Texas costing approximately $14 billion, underscoring the scale of capital being poured into AI compute infrastructure where GPU utilization efficiency directly determines return on investment. The pressure to extract more throughput per GPU, rather than simply adding more cards, is what makes MPS-style optimizations commercially significant for streaming and speech workloads. The competitive landscape for AI inference hardware is intensifying, with alternatives to NVIDIA's GPU architecture entering public markets. Cerebras Systems, known for its wafer-scale engine design, filed for an IPO with a reported $10 billion contract from OpenAI forming a cornerstone of its growth narrative. The company's architecture targets massive parallelism with lower latency than traditional GPU clusters, and its partnership with Amazon Web Services signals that hyperscalers are actively diversifying their inference silicon options. For streaming platforms evaluating ASR pipelines on AWS, the emergence of credible alternatives to NVIDIA GPUs could eventually pressure pricing and accelerate adoption of efficiency techniques like MPS in the interim. On the deployment side, voice AI workloads are increasingly being packaged as managed cloud services that abstract GPU orchestration from end users. Deepgram, a speech-to-text provider, integrated its real-time STT and TTS models as SageMaker-ready endpoints available through AWS Marketplace, running inference inside customer VPCs with sub-300 ms end-to-end latency using its Flux and Nova models. The deployment inherits the customer's IAM, VPC, KMS, and CloudWatch security posture, demonstrating the pattern that AWS and NVIDIA's MPS guide targets: keeping sensitive audio data within controlled boundaries while maximizing GPU throughput. For streaming platforms building live captioning, contact center transcription, or voice-agent features, the convergence of MPS-level GPU efficiency and managed endpoint packaging represents the operational model toward which the industry is converging.
Read full article at aws.amazon.com
Enjoy our coverage?
Add StreamingMeme as a preferred source on Google to see more of our streaming news at the top of your Search results.
Add as preferred source