Multiverse Computing raises $570M to shrink AI models for the edge
Multiverse Computing has raised $570 million in Series C funding to advance its CompactifAI technology, which utilizes tensor networks to compress AI models for efficient edge and CPU-based inference. The company aims to apply this technology to local on-premises AI deployment for enterprise customers, reducing the need for cloud-based hyperscaler infrastructure.
Key Takeaways
- Series C funding co-led by Forgepoint Capital, BNPP SIVF, and Bullhound Capital at a $1.7 billion valuation
- CompactifAI technology reduces model size by 80% to 95% with minimal accuracy loss using tensor networks
- Compressed Meta Llama 3.3 70B model successfully runs on Intel Xeon 6 CPUs, cutting disk size from 130GB to 65GB
- Software stack includes a real-time router that automatically directs workloads to either local hardware or the cloud
- Strategic expansion targets AI gigafactory infrastructure and regional growth in East Asia, the Middle East, and North America
Why It Matters
Multiverse breaks the assumption that high-performance AI requires massive GPU clusters. By enabling 70B-parameter models to run on standard CPUs and edge hardware, the startup shifts the economics of streaming and enterprise AI from expensive Opex-heavy cloud reliance to local, controlled infrastructure. For the streaming ecosystem, this facilitates low-latency, privacy-compliant AI applications—such as real-time content moderation or hyper-personalized UI—directly on consumer devices. Watch for Intel's and HP's strategic involvement to accelerate the adoption of 'sovereign AI' that operates entirely behind enterprise firewalls.
Additional Context
The push for local inference comes as enterprises seek to mitigate the volatile costs of cloud-based AI. Per The Next Web (July 2026), the AI inference market is projected to reach over $255 billion by the early 2030s, shifting the industry focus from model training to efficient deployment. Multiverse has reported significant commercial traction in this transition, noting that its annualized revenue grew 10x following its June 2025 Series B round, with Q1 2026 sales recording a 96-fold year-over-year increase.
Hardware partnerships are central to this efficiency play. According to Multiverse's own technical benchmarks (July 2026), its compressed version of Llama 3.3 70B achieves a 94.1% increase in total token throughput on Intel Xeon 6 processors compared to uncompressed baselines. This performance gain is driven by Intel’s Advanced Matrix Extensions (AMX), allowing models that typically require high-end GPUs to operate on general-purpose data center CPUs. The company's expansion into 'AI gigafactory' infrastructure suggests a shift toward vertically integrated software-hardware stacks for private enterprise clouds.
Sector-wide, the concentration of capital in AI infrastructure remains dominant. Per Fundraise Insider (July 2026), AI captured nearly 50% of all global startup funding in 2025, totaling $202.3 billion. Multiverse’s $570 million round participates in a broader trend of mega-rounds for companies addressing the 'scalability crisis'—the point where model size outpaces the energy and cost limits of traditional data centers. The inclusion of strategic institutional investors like Qatar Development Bank and Orange Ventures indicates that sovereign control over AI processing has become a priority for national and telecommunications infrastructure providers.
Read full article at siliconangle.com
Enjoy our coverage?
Add StreamingMeme as a preferred source on Google to see more of our streaming news at the top of your Search results.
Add as preferred source