KubeCon North America 2026 adds AI inference track for production workloads
KubeCon + CloudNativeCon North America 2026, scheduled for November 9-12 in Salt Lake City, will feature a new AI Inference + Agentic track. The event will focus on operationalizing AI infrastructure, including model serving and GPU scheduling, alongside ongoing themes of platform engineering and security.
Key Takeaways
- New AI Inference + Agentic track focuses on model serving, GPU scheduling, and inference performance
- Featured open-source projects include vLLM, KServe, Ray, and OpenTelemetry for AI observability
- Platform engineering sessions will highlight Backstage, Argo, and GitOps for enterprise internal platforms
- VMware will demonstrate convergence between virtualization and containers via VMware vSphere Kubernetes Service
Why It Matters
The introduction of a dedicated inference track signals a shift from model training to the complex operational realities of serving AI at scale. For streaming and cloud-native providers, this transition requires sophisticated GPU scheduling and observability to manage high-concurrency workloads efficiently. As Kubernetes becomes the standard for production AI, the convergence of virtualization and container orchestration—exemplified by VMware’s latest integrations—will likely define the next phase of private cloud architecture. Watch for specific performance benchmarks from the vLLM and KServe projects during the November sessions to gauge the maturity of open-source inference stacks.
Additional Context
The Cloud Native Computing Foundation has been aggressively expanding its portfolio of AI-focused projects to meet demand from enterprises operationalizing inference workloads on Kubernetes. In March 2026, CNCF announced that vLLM inference engine had been accepted as a sandbox project, joining KServe and Ray as part of a growing inference ecosystem. The foundation's 2025 annual survey found that 68% of respondents reported running AI or ML workloads in production on Kubernetes, up from 52% the prior year, underscoring why KubeCon's programming has shifted toward inference operations rather than model training. VMware, now under Broadcom, has positioned its vSphere Kubernetes Service as a bridge for enterprises that want to run GPU-accelerated inference without abandoning their existing virtualization investments. On the business and standards side, Broadcom completed its acquisition of VMware in November 2023 and has since restructured the product line around VMware Cloud Foundation as the primary SKU. In June 2026, Broadcom announced that VMware Cloud Foundation 9 would include native GPU scheduling for AI inference workloads, directly competing with hyperscaler-managed inference services. The move targets enterprises that want to run vLLM or KServe on-premises with the same orchestration semantics they use in public cloud. Meanwhile, the Linux Foundation's OpenTelemetry project has been extending its semantic conventions to cover GPU utilization metrics, which were formally proposed in a specification pull request in May 2026, giving platform teams a standardized way to observe inference performance across heterogeneous hardware. Technical benchmarks for the inference stack are maturing rapidly. In April 2026, Anyscale published results showing that Ray Serve combined with vLLM achieved 2.3x higher throughput per GPU compared to standalone vLLM deployments on Llama 3 70B serving, with p99 latency remaining under 400 milliseconds at 512 concurrent requests. KServe, which provides model serving abstractions on top of Kubernetes, released version 0.14 in July 2026 with support for speculative decoding and multi-LoRA adapter routing, features that directly address the multi-tenant inference challenges streaming platforms face when serving personalized recommendation models. These advances suggest that the sessions at KubeCon North America in November will likely focus less on whether Kubernetes can handle inference and more on how to optimize cost and latency at scale. For enterprises looking to scale these deployments, with the latest updates to its infrastructure stack. As gains traction, these operational patterns will become critical for global media delivery.
Read full article at virtualizationreview.com
Enjoy our coverage?
Add StreamingMeme as a preferred source on Google to see more of our streaming news at the top of your Search results.
Add as preferred source