IBM secures $240M deal providing Nvidia Blackwell systems to Together AI
IBM has entered into a $240 million multi-year infrastructure agreement with Together AI to provide Nvidia HGX B300-based systems. This hardware will power Together AI's public cloud platform, specifically for training and inferencing open-source AI models.
Key Takeaways
- IBM will deploy a large-scale cluster of Nvidia HGX B300 motherboards, each featuring eight Blackwell Ultra chips and Spectrum-X Ethernet networking.
- Together AI currently processes 400 trillion tokens per month, serving customers such as Cursor, Cognition, and Decagon.
- The new infrastructure is designed to deliver 30 times more AI factory output than previous generations for high-throughput inference workloads.
- Hardware for the cluster is scheduled to become available through IBM Cloud in the first half of 2027.
Why It Matters
The deal signals a significant shift in the cloud market as IBM positions itself as a primary infrastructure provider for "neoclouds" like Together AI. By focusing on production-grade inference rather than just model training, IBM and Together AI are targeting enterprises seeking to scale open-source models at a lower cost than proprietary APIs. This partnership addresses the increasing demand for hardware that supports the high-throughput, low-latency requirements of media-generation and agentic AI models. Streaming and video platforms should monitor whether this expansion reduces the high per-token costs of AI-driven personalization and automated content tagging.
Additional Context
The IBM agreement follows a period of massive capital injection into the neocloud sector. In July 2026, Together AI closed an $800 million Series C funding round at an $8.3 billion valuation, led by Aramco Ventures with participation from Nvidia and Salesforce Ventures. According to TechCrunch, the company's annual bookings recently surpassed $1.15 billion, driven by developers seeking to avoid the high margins of closed-model providers. This trend is mirrored by competitors like Groq, which raised $650 million in June 2026 to scale its own inference cloud after licensing chip technology to Nvidia in late 2025. Industry spending patterns are shifting rapidly toward these dedicated inference environments. Per Gartner reporting from August 2026, global spending on AI-optimized IaaS is projected to reach $42 billion this year, a 96% increase. For the first time, spending on inference ($23.3 billion) is expected to surpass training ($19 billion), as enterprises move beyond experimental pilots into full-scale production. Gartner analysts note that the rise of "agentic AI"—which executes multi-step autonomous tasks—is amplifying compute intensity, making dedicated clusters like the one IBM is building a requirement for enterprise scalability. IBM's strategic pivot toward Nvidia-backed hardware also expands a collaboration first detailed in March 2026. At that time, IBM announced plans to integrate Nvidia's Blackwell Ultra GPUs into its Red Hat AI Factory stack and Watsonx.data platform. This recent $240 million deal with Together AI represents the first concrete, large-scale deployment of that integrated roadmap on IBM Cloud, specifically targeting the open-weight model market including Kimi, Mistral, and DeepSeek.
Read full article at siliconangle.com
Get this in your inbox → Subscribe
Enjoy our coverage?
Add StreamingMeme as a preferred source on Google to see more of our streaming news at the top of your Search results.
Add as preferred source