IBM and Together AI sign $240M deal for NVIDIA inference infrastructure
IBM and Together AI have entered into a $240 million multi-year agreement to build a large-scale inference cluster on IBM Cloud. The infrastructure, powered by NVIDIA HGX B300 systems, is designed to support enterprise-grade, open-source AI model inference starting in Q1 2027.
Key Takeaways
- The agreement includes a $240 million commitment over multiple years to scale Together AI's dedicated inference capacity.
- IBM Cloud will deploy NVIDIA HGX B300 systems featuring Spectrum-X Ethernet networking, with target availability in Q1 2027.
- Infrastructure performance is expected to deliver 30 times more AI factory output than previous generations.
- Together AI currently processes 400 trillion tokens per month and recently closed an $800 million Series C funding round.
Why It Matters
This partnership marks a shift toward specialized AI inference infrastructure, prioritizing the economical scaling of open-source models over closed-source APIs. For the streaming and media ecosystem, the deployment of NVIDIA B300 systems suggests a move toward agentic AI workflows and complex reasoning that require high memory capacity for real-time video processing and personalization. As enterprise demand moves from model training to production-scale inference, the availability of 30x higher throughput at a lower token cost will be the benchmark for technical feasibility. Watch for whether this capacity helps Together AI reach its projected $1 billion annualized revenue target by early 2027.
Additional Context
The expansion follows a massive growth period for Together AI, which reported a 13,000x increase in monthly token volume over a single year, reaching 400 trillion tokens by June 2026 per KuCoin. To sustain this trajectory, the company raised $800 million in July 2026 in a round led by Aramco Ventures, valuing the startup at $8.3 billion, according to Sacra. This valuation represents a 2.5x increase in just 17 months, driven by the company's focus on providing a modular stack for developers who prefer open-source flexibility over the pricing of proprietary models.
The hardware at the center of this deal, the NVIDIA Blackwell Ultra B300, is specifically engineered for the high-concurrency demands of agentic AI. Launched to meet the needs of models that require extensive multi-step reasoning, the B300 features 288 GB of HBM3e memory—a 50% increase over the standard B200 model, according to NVIDIA specifications from mid-2026. This additional memory allows 300B+ parameter models to run entirely on a single GPU without paging to CPU memory, which dramatically reduces latency for real-time enterprise services.
IBM's role as the infrastructure provider signals its intent to recapture enterprise cloud market share by positioning itself as a primary host for high-end NVIDIA Blackwell clusters. In March 2026, per IBM, the company expanded its broader collaboration with NVIDIA to include GPU-native data analytics and unstructured data extraction. By providing Together AI with the first dedicated large-scale B300 cluster for inference, IBM is competing directly with hyperscalers like Azure and AWS for the next generation of AI-native cloud workloads.
Read full article at newsroom.ibm.com
Enjoy our coverage?
Add StreamingMeme as a preferred source on Google to see more of our streaming news at the top of your Search results.
Add as preferred source