IBM and Together AI sign $240M deal for NVIDIA Blackwell systems
IBM has entered into a $240 million multi-year agreement to supply Together AI with NVIDIA HGX B300 infrastructure, scheduled for deployment in the first half of 2027. The deal scales the cloud operator's capability to process inference workloads, including those optimized for media generation and enterprise-grade AI models.
Key Takeaways
- Together AI will utilize NVIDIA Blackwell Ultra chips to process approximately 400 trillion tokens per month for its customers.
- The infrastructure includes HGX B300 motherboards featuring 2.3TB of combined HBM3e memory and NVIDIA Spectrum-X Ethernet networking.
- Deployment is scheduled for the first half of 2027, marking the first dedicated Blackwell-based inference cluster on IBM Cloud.
- The agreement follows Together AI's recent $800 million Series C funding round, which valued the startup at $8.3 billion.
Why It Matters
The $240 million commitment signals a shift in the streaming and media ecosystem toward specialized neocloud providers for high-volume inference. By utilizing the NVIDIA HGX B300, Together AI targets the high-compute demands of media generation models, which require massive memory bandwidth to remain cost-effective. For the broader industry, this deal validates the enterprise transition from closed-source APIs to open-weight AI models for production-grade workloads. It also solidifies IBM’s position as a critical infrastructure partner for high-growth AI startups challenging legacy cloud giants. Watch for whether this deployment triggers a price war in per-token inference costs as more Blackwell-based clusters come online in early 2027.
Additional Context
The partnership between IBM and Together AI aligns with a broader industry trend where inference workloads are projected to account for two-thirds of all AI compute by late 2026. Per industry reporting from July 2026, the global AI inference market is expected to reach $117.8 billion this year, driven by the rapid adoption of multimodal generation and long-context reasoning. These advanced tasks often require up to 100 times the compute of simple queries, making high-memory systems like the Blackwell Ultra essential for maintaining latency requirements in production environments. Together AI’s $800 million Series C round in July 2026, led by Aramco Ventures with participation from NVIDIA, underscored the rising valuation of the 'neocloud' category. According to TechCrunch, the startup’s annual bookings surpassed $1.15 billion as enterprises sought to reduce costs by up to 60 times compared to closed-model providers. This capital influx has sparked an infrastructure arms race among specialized providers; for example, rival TensorWave raised $350 million in June 2026 to scale its own inference-focused clusters. Technically, the HGX B300 systems at the heart of the IBM deal represent a significant jump in density. Per NVIDIA’s 2026 specifications, the Blackwell Ultra architecture provides 15 petaflops of dense FP4 performance per chip, a 1.5-times improvement over the standard B200. These systems are specifically designed to handle trillion-parameter models with long-context requirements, using 288GB of HBM3e memory per GPU to keep large models entirely in-memory, thereby eliminating the throughput bottlenecks common in older hardware generations.
Read full article at siliconangle.com
Enjoy our coverage?
Add StreamingMeme as a preferred source on Google to see more of our streaming news at the top of your Search results.
Add as preferred source