TileRT achieves 500 tokens per second on NVIDIA B200 GPUs
The research platform InferenceX analyzed TileRT software, which optimizes NVIDIA B200 GPUs for ultra-high interactivity inference by compiling decode graphs into persistent kernels. The technology enables GPUs to reach 500 tokens/s/user, providing a software-based alternative to specialized inference ASICs for latency-sensitive applications like real-time AI assistants.
Key Takeaways
- TileRT reached 340 tokens/s/user on an 8-GPU B200 node, roughly 1.9x faster than results from the 72-GPU GB300 NVL72 rack.
- The software uses persistent engine kernels to eliminate the latency overhead caused by traditional GPU kernel setup and teardown.
- Xiaomi has already deployed the TileRT decode engine in production for its MiMo V2.5 Pro UltraSpeed AI model.
- Performance gains are currently limited to a batch size of 1, optimized specifically for ultra-low-latency interaction rather than bulk throughput.
- Integration with vLLM allows a disaggregated approach where TileRT handles sensitive decode tasks while vLLM manages compute-intensive prefill.
Why It Matters
TileRT reframes the inference market by allowing streaming and AI providers to provision 'speed tiers' from existing GPU fleets rather than purchasing dedicated ASICs from startups like Groq or Cerebras. This software-driven approach provides critical flexibility, enabling operators to rebalance capacity between high-throughput batching and ultra-interactive real-time modes via configuration rather than hardware re-cabling. For B2B streaming services integrating AI voice assistants, this reduces the total cost of ownership (TCO) while meeting the sub-millisecond latency requirements of full-duplex conversational models. Watch for TileRT’s upcoming 'AgentX' benchmarks to see if these interactivity gains hold up during long-context, multi-turn reasoning tasks exceeding 140,000 tokens.
Additional Context
The push for ultra-high interactivity follows the launch of high-stakes AI voice products such as OpenAI’s GPT-Live. Per OpenAI in August 2026, real-time conversational systems now utilize full-duplex architectures that listen and speak simultaneously, making even minor token generation delays perceptible to users. This shift has placed immense pressure on inference providers to lower Time Per Output Token (TPOT) to levels previously thought impossible on general-purpose GPUs. While NVIDIA dominated the high-throughput era, the rise of agentic AI requires a specialized 'interactivity' focus that has historically favored the deterministic scheduling of dedicated dataflow chips.
Market adoption of software-based optimization is accelerating as an alternative to purpose-built hardware. Xiaomi’s recent release of the MiMo V2.5 Pro UltraSpeed mode in June 2026 demonstrated that model-system co-design can break the 1,000 tokens-per-second barrier for specific use cases. Per Xiaomi, this high-speed mode was offered at 3x the standard API cost, proving a viable commercial path for low-latency services. This strategy leverages existing commodity GPU infrastructure, allowing companies to avoid the lead times and capital risk associated with scaling custom ASIC clusters.
However, the competitive landscape remains fragmented. While TileRT improves standard GPU performance, specialized vendors like Cerebras and Groq continue to push the absolute ceiling of performance. Per SemiAnalysis and MLPerf reporting from early 2026, dedicated hardware often still leads in raw throughput for massive dense models that cannot be easily optimized through software kernel persistent alone. The industry is currently split between those prioritizing the fungibility of a unified GPU fleet and those seeking the extreme efficiency of hardwired transformer accelerators like the Etched Sohu or Taalas HC1.
Read full article at newsletter.semianalysis.com
Enjoy our coverage?
Add StreamingMeme as a preferred source on Google to see more of our streaming news at the top of your Search results.
Add as preferred source