Sharon AI taps Rafay Systems to manage 150,000 GPUs globally
Sharon AI has signed a five-year agreement with Rafay Systems to utilize its platform for orchestrating up to 150,000 GPUs across global AI Factory environments. The partnership aims to standardize infrastructure management, provisioning, and governance for AI workloads including training and inference.
Key Takeaways
- The five-year deal supports the orchestration of up to 150,000 GPUs across multiple global locations.
- Rafay AI infrastructure platform provides a unified control plane for provisioning, monitoring, and tenant isolation.
- Sharon AI is transitioning from managing individual clusters to a standardized operating model for its expanding Asia-Pacific footprint.
- The partnership covers infrastructure lifecycle management for both model training and inference workloads.
Why It Matters
This partnership highlights the shift from raw compute acquisition to sophisticated operational management in the AI infrastructure sector. By implementing a centralized orchestration layer, Sharon AI addresses the complexity of managing massive GPU fleets while maintaining secure multi-tenant isolation. For the streaming and AI ecosystem, this reflects a growing reliance on specialized software to maximize the utilization of scarce hardware resources. As neocloud providers compete with hyperscalers, the ability to automate service delivery and governance becomes a critical differentiator. Watch for Sharon AI to report specific utilization improvements and deployment speed metrics as they integrate Rafay software across their Asia-Pacific data centers.
Additional Context
Rafay Systems has been building momentum in the GPU orchestration space beyond its Sharon AI partnership. In March 2026, Rafay Systems announced a $100 million Series C funding round led by Tiger Global, valuing the company at $1 billion and bringing total funding to $175 million. The round was earmarked for expanding its AI infrastructure platform into new geographies and deepening support for multi-cloud GPU fleet management. That capital injection positions Rafay to compete directly with orchestration offerings from hyperscalers and other neocloud software vendors targeting the same AI Factory buildout wave that Sharon AI is riding.
The neocloud market itself is drawing significant investor and operator attention. In July 2026, CoreWeave reported revenue of $1.2 billion for its second quarter, surpassing analyst expectations and signaling sustained demand for GPU-as-a-service capacity from AI model developers. CoreWeave's scale validates the neocloud model that Sharon AI is pursuing at a smaller but rapidly growing footprint. Meanwhile, Lambda Labs secured a $1.5 billion debt facility in May 2026 to finance GPU cluster expansion, underscoring how capital-intensive GPU infrastructure has become and why orchestration software that improves utilization rates carries outsized economic value for operators managing tens of thousands of accelerators.
On the technical side, Rafay Systems has differentiated its platform through support for both Kubernetes-native and virtual machine workloads within a single control plane. In a June 2026 benchmark published by The New Stack, Rafay's platform demonstrated a 40% reduction in GPU provisioning time compared to manual cluster configuration, a metric that directly affects the economics of AI Factory operations. For Sharon AI, which plans to scale to 150,000 GPUs across multiple regions, even marginal gains in provisioning speed and utilization translate into meaningful cost savings. The platform also integrates policy-based governance for multi-tenant isolation, a requirement that becomes critical as neocloud providers like Sharon AI serve enterprise customers with strict data residency and compliance needs.
Read full article at datacenternews.asia
Enjoy our coverage?
Add StreamingMeme as a preferred source on Google to see more of our streaming news at the top of your Search results.
Add as preferred source