SemiAnalysis open sources $3M AgentX coding inference benchmark for agentic workloads
SemiAnalysis has released AgentX 1.0, an open-source benchmark designed to measure multi-turn agentic coding inference performance at 1 million context length. The benchmark provides data on hardware performance from Nvidia and AMD, highlighting the importance of system-level optimizations like KV cache management in production agentic workloads.
Key Takeaways
- Nvidia B300 vLLM achieved a 91% HBM cache hit rate under a load of 384 concurrent agentic traces
- AMD MI355X performance per dollar surpassed B200 vLLM in specific Kimi K3 model scenarios using the ATOM engine
- The benchmark has already driven over 70 upstream pull requests for optimizations in vLLM, SGLang, and TensorRT-LLM
- Nvidia B300 FP4 demonstrated 12x better performance per dollar compared to H100 when running Qwen3.5 397B
Why It Matters
The release of AgentX signals a shift from measuring raw chip kernels to evaluating entire inference systems, where KV cache management and routing affinity determine production viability. For streaming and AI infrastructure providers, this data proves that HBM capacity and system-level software optimizations are now as critical as peak FLOPs for long-context applications. While Nvidia currently leads in high-throughput scenarios for models like MiniMax M3, AMD is closing the gap in specific frontier model architectures through its ATOM engine. Watch for the upcoming AgentX update in three weeks to see if AMD can successfully upstream these performance gains into the broader vLLM ecosystem.
Additional Context
SemiAnalysis has positioned itself as an independent arbiter of AI infrastructure performance, and AgentX extends that role into agentic workloads specifically. The firm's InferenceX benchmark series has previously evaluated large language model serving across Nvidia's H100 and B200 GPUs, and SemiAnalysis published its InferenceXv2 results in early 2025 comparing H100, H200, and B200 on Llama-class models, establishing a methodology that AgentX now builds upon by adding multi-turn context and sub-agent orchestration patterns. The release also arrives as Nvidia prepares its GB300 NVL72 rack-scale systems for volume shipment, with Nvidia confirming at Computex 2025 that GB300 would begin shipping to hyperscalers in the second half of 2025 targeting inference-heavy workloads that demand the kind of long-context KV cache management AgentX measures.
The competitive dynamics between Nvidia and AMD in inference hardware have intensified through 2025 and into 2026. AMD launched its MI355X accelerator at Advancing AI 2025, positioning the chip as a direct competitor to Nvidia's B200 for both training and inference with 288 GB of HBM3e memory, a memory capacity advantage that becomes relevant for the million-token context windows AgentX tests. Meanwhile, the open-source inference software stack that both vendors target has consolidated around vLLM and SGLang. vLLM crossed 50,000 GitHub stars in mid-2025 and added support for speculative decoding and prefix caching optimizations that directly affect the KV cache reuse patterns AgentX quantifies. SGLang, developed at UC Berkeley and now backed by LMSYS, released its RadixAttention mechanism for automatic prefix caching in production serving, a technique that AgentX's benchmark design explicitly stress-tests through repeated sub-agent calls sharing context.
Independent benchmarking of inference hardware has become a growing field as enterprises move from production deployment to enterprise AI agent security. MLPerf, the industry-standard benchmark consortium, added inference scenarios for Llama 2 70B and Mixtral 8x7B in its v4.1 results published in late 2024, but those tests use fixed prompt lengths and do not capture the multi-turn agentic patterns that AgentX introduces. The gap between synthetic benchmarks and production reality has been a recurring theme; Anthropic reported in its engineering blog that real Claude API traffic exhibits bursty concurrency and context reuse patterns that differ significantly from uniform-load benchmarks, validating the design philosophy behind AgentX's workload modeling. For streaming infrastructure providers evaluating GPU clusters for AI-powered content recommendation, transcoding orchestration, or conversational interfaces, AgentX provides the first standardized dataset that reflects how those systems will actually behave under production agentic loads rather than isolated single-request latency tests.
Read full article at newsletter.semianalysis.com
Enjoy our coverage?
Add StreamingMeme as a preferred source on Google to see more of our streaming news at the top of your Search results.
Add as preferred source