Sakana AI Fugu orchestrators launch to slash multi-agent reasoning costs
Sakana AI has launched Fugu Max and Fugu Ultra v2, two learned orchestrators that route queries across a pool of models via a single API. Fugu Max is designed for cost-efficiency, while Fugu Ultra v2 focuses on high-performance reasoning and software engineering tasks.
Key Takeaways
- Fugu Max pricing is set at $2 per 1M input tokens and $6 per 1M output tokens, targeting the cost-performance Pareto frontier.
- Fugu Ultra v2 achieved a 48.3 score on Chartography visual reasoning, significantly outperforming Opus 5 and Fable 5.
- The orchestration architecture uses TRINITY and Conductor frameworks to assign Thinker, Worker, and Verifier roles across model turns.
- Sakana AI integrated NVIDIA Nemotron into the Fugu Max model pool through a strategic collaboration with NVIDIA.
Why It Matters
The release of these orchestrators signals a shift toward model-agnostic infrastructure where performance is decoupled from single-vendor dependencies. By routing tasks to the leanest capable agent, Sakana AI allows streaming engineers to optimize high-volume metadata and software tasks without the premium cost of frontier models. This approach mitigates risks associated with API revocations or sudden service cutoffs from major providers. As the ecosystem moves toward multi-agent workflows, the industry should monitor whether this orchestration layer can maintain its 60% cost advantage as proprietary models like GPT-6-Astra enter the broader market.
Additional Context
Nokia has emerged as the most aggressive vendor in deploying agentic AI across its full network stack, directly competing with the multi-agent orchestration approach Sakana AI is bringing to the model-routing layer. In June 2026, Nokia launched an agentic AI framework built into its Network Services Platform for IP network operations, marking its third agentic product announcement in four weeks. The framework lets carriers deploy AI agents that make decisions from real-time network data and take autonomous action within operator-defined guardrails, a design philosophy that mirrors Sakana AI's learned routing across model pools.
The business architecture around Nokia's agentic push is equally significant. At DTW Ignite in June 2026, Nokia announced partnerships with AWS and Databricks to build the data, cloud, and control layers for autonomous networks, positioning its Autonomous Network Fabric as an operating system spanning radio, core, transport, and service domains. Nokia claims operators using its autonomous networks portfolio are achieving automation rates above 90 percent, service delivery times under four hours, and up to 85 percent reduction in slice rollout time. Meanwhile, Ericsson launched its AI in RAN commercial software subscription on June 11, claiming up to 20 percent higher downlink throughput across more than 15 live deployments using existing baseband silicon, while Verizon disclosed that its 60,000-site vRAN is now applying agentic AI to configuration changes and network optimization.
The technical divergence between vendors on how agents interact with hardware has direct implications for orchestration design. Ericsson and Nokia are diverging sharply on AI-RAN architecture, with Nokia running all Layer 1 functions on Nvidia GPUs via CUDA while Ericsson confines only the FEC function to the GPU and runs remaining L1 software on CPUs. Nokia is also embedding agentic AI deeper into its mobile core, where the company claims AI-driven paging reduces call setup time from roughly 10 seconds to one or two seconds by using smaller models collocated with network functions for edge inferencing. These deployment patterns validate the broader thesis behind Sakana AI's orchestrators: that intelligent routing across heterogeneous compute and model resources, rather than reliance on a single monolithic system, delivers measurable efficiency gains at scale.
Read full article at marktechpost.com
Enjoy our coverage?
Add StreamingMeme as a preferred source on Google to see more of our streaming news at the top of your Search results.
Add as preferred source