IBM and Together AI have entered into a $240 million multi-year agreement to build a large-scale inference cluster on IBM Cloud. The infrastructure, powered by NVIDIA HGX B300 systems, is designed to support enterprise-grade, open-source AI model inference starting in Q1 2027.
This partnership marks a shift toward specialized AI inference infrastructure, prioritizing the economical scaling of open-source models over closed-source APIs. For the streaming and media ecosystem, the deployment of NVIDIA B300 systems suggests a move toward agentic AI workflows and complex reasoning that require high memory capacity for real-time video processing and personalization. As enterprise demand moves from model training to production-scale inference, the availability of 30x higher throughput at a lower token cost will be the benchmark for technical feasibility. Watch for whether this capacity helps Together AI reach its projected $1 billion annualized revenue target by early 2027.
The expansion follows a massive growth period for Together AI, which reported a 13,000x increase in monthly token volume over a single year, reaching 400 trillion tokens by June 2026 per KuCoin. To sustain this trajectory, the company raised $800 million in July 2026 in a round led by Aramco Ventures, valuing the startup at $8.3 billion, according to Sacra. This valuation represents a 2.5x increase in just 17 months, driven by the company's focus on providing a modular stack for developers who prefer open-source flexibility over the pricing of proprietary models.
The hardware at the center of this deal, the NVIDIA Blackwell Ultra B300, is specifically engineered for the high-concurrency demands of agentic AI. Launched to meet the needs of models that require extensive multi-step reasoning, the B300 features 288 GB of HBM3e memory—a 50% increase over the standard B200 model, according to NVIDIA specifications from mid-2026. This additional memory allows 300B+ parameter models to run entirely on a single GPU without paging to CPU memory, which dramatically reduces latency for real-time enterprise services.
IBM's role as the infrastructure provider signals its intent to recapture enterprise cloud market share by positioning itself as a primary host for high-end NVIDIA Blackwell clusters. In March 2026, per IBM, the company expanded its broader collaboration with NVIDIA to include GPU-native data analytics and unstructured data extraction. By providing Together AI with the first dedicated large-scale B300 cluster for inference, IBM is competing directly with hyperscalers like Azure and AWS for the next generation of AI-native cloud workloads.
Add StreamingMeme as a preferred source on Google to see more of our streaming news at the top of your Search results.
Add as preferred source