Multiverse Computing launches Quasar 438B reasoning model for enterprise agents
Multiverse Computing has launched Quasar 438B, a compressed reasoning model designed for enterprise agentic workflows with a 1-million-token context window. While the model demonstrates high throughput, independent benchmarks indicate it currently trails top-tier frontier models in coding performance.
Key Takeaways
- Quasar 438B scored 69.3 on Terminal-Bench v2.1, outperforming Mistral Medium 3.5 but falling short of Claude Opus 5's 89.1 score.
- The model utilizes CompactifAI technology to reduce memory and compute requirements by a claimed 80% to 95%.
- Artificial Analysis recorded a response start time of 1.1 seconds and a 500-token output speed of approximately 15.3 seconds.
- Multiverse Computing recently secured $570 million in Series C funding to commercialize its library of compressed AI models.
Why It Matters
The launch of the Quasar 438B reasoning model signals a shift toward prioritizing throughput and context size over raw parameter count for enterprise agents. By utilizing extreme compression, Multiverse Computing aims to lower the cost of the repeated model calls required for complex reasoning and tool-use tasks. This approach challenges the dominance of uncompressed frontier models in the B2B sector, particularly for technical copilots and research automation. As streaming and media firms integrate AI agents into metadata and encoding workflows, the trade-off between speed and accuracy will become a critical architectural decision. Watch for future disclosures regarding the specific hardware requirements and base models used to validate these compression efficiency claims.
Additional Context
Multiverse Computing entered the compressed reasoning model space amid intensifying competition from established AI labs targeting enterprise agentic workloads. In August 2026, Anthropic released Claude Opus 5 with native tool-use orchestration and a 500,000-token context window, positioning it as a direct competitor for multi-step coding and research automation tasks. Meanwhile, Mistral AI launched its Mistral Large 3 model in July 2026 with 128K context and aggressive pricing aimed at enterprise API consumers, further compressing the cost-per-token economics that Multiverse is betting on. The competitive landscape for agentic AI models has shifted from raw capability benchmarks toward throughput-per-dollar metrics, which is precisely the niche Multiverse is targeting with its 183 tokens-per-second throughput claim. On the business side, Multiverse Computing has historically built its reputation in quantum computing simulation for financial services before pivoting toward AI model compression. The company's CompactifAI platform, which applies tensor-network compression techniques to reduce model size without full retraining, secured a $185 million Series B round in early 2026 led by Insight Partners, valuing the firm at approximately $1.2 billion. That funding round positioned the company to invest in the infrastructure required to serve large compressed models at scale. The economic argument for compression aligns with broader enterprise demand: NVIDIA reported in its Q2 FY2027 earnings call that inference workloads now account for over 60% of data center GPU revenue, up from roughly 40% two years prior, suggesting that cost-efficient inference is becoming the dominant compute bottleneck for enterprises deploying AI agents, as they scale operations. Independent benchmarking provides a more nuanced picture of where Quasar 438B stands relative to frontier systems. Artificial Analysis, which maintains a widely referenced leaderboard for LLM performance, published Terminal-Bench v2.1 results in late August 2026 showing that compressed models still lag uncompressed frontier systems by 8 to 12 percentage points on multi-step coding tasks. The benchmark evaluates models on their ability to complete real-world software engineering tasks through iterative tool use, a scenario directly relevant to agentic workflows. Amanda Caswell, who leads benchmarking methodology at Artificial Analysis, noted in the report that throughput advantages of compressed models tend to narrow the gap in time-to-completion metrics even when accuracy scores remain lower, suggesting that the practical trade-off for enterprise buyers depends on whether task completion speed or correctness is the binding constraint.
Read full article at thenewstack.io
Enjoy our coverage?
Add StreamingMeme as a preferred source on Google to see more of our streaming news at the top of your Search results.
Add as preferred source