Cisco tackles agentic AI economics with Antares models and Cloud Control
Cisco Systems President Jeetu Patel discussed the impact of agentic AI on enterprise infrastructure at the AMD Advancing AI event. Cisco is addressing the resulting token-consumption demand by introducing the Antares open-weight model family and the Cisco Cloud Control management plane to monitor hybrid inference workloads.
Key Takeaways
- Antares-350M and Antares-1B are specialized open-weight models designed for vulnerability localization and local repository exploration.
- Cisco Cloud Control provides a unified management plane to monitor hybrid inference workloads across cloud, private data centers, and endpoint devices.
- Internal testing showed a 500-entry evaluation run using Antares cost less than $1 and completed in approximately 15 minutes on a single GPU.
- The 1B Antares model achieves a 0.209 File F1 score on the VLoc Bench, outperforming general-purpose models at a fraction of the size.
Why It Matters
The transition from human-prompted chatbots to autonomous agents creates a shift toward continuous inference, making 'tokenomics' a central infrastructure challenge for streaming and enterprise stacks. Cisco’s strategy emphasizes decentralizing compute, allowing high-volume tasks like security triage to run locally on specialized models while reserving expensive frontier models for complex reasoning. This hybrid approach addresses the supply shortage for AI compute by reducing reliance on high-end cloud instances. For the streaming industry, which relies on large-scale automated code and infrastructure management, this signals a move toward more granular, cost-managed AI observability. Watch for Cisco's release of the 3B parameter Antares model as a benchmark for local-first enterprise inference performance.
Additional Context
The expansion into local, agentic operations comes as enterprise AI spending undergoes a significant structural shift. Per McKinsey & Company (July 2026), 93% of enterprises now report exceeding their AI budgets despite a 99% collapse in the unit cost of tokens over the last two years. This divergence is largely attributed to the high volume of 'response refinement' in agentic loops, which currently accounts for 60% of total agent cost. Cisco’s focus on 'tokenomics' directly targets this volume problem by offering visibility into which agents are consuming resources excessively across distributed hybrid environments. Cisco’s partnership with AMD, highlighted at the Advancing AI 2026 event, also reflects an industry-wide pivot toward on-device inference. Per Network World (July 2026), the joint architecture pairs AMD’s Ryzen AI Halo hardware—capable of running 300-billion-parameter models locally—with Cisco’s Splunk Agent Observability and Cloud Control software. This collaboration aims to provide a secure 'harness' for autonomous agents, ensuring that even when agents communicate with other agents across a corporate network, their actions remain governed and observable. Furthermore, the release of the Antares family on July 21, 2026, includes the open-sourcing of the Vulnerability Localization Benchmark (VLoc Bench). Per TechRepublic (July 2026), the benchmark consists of 500 tasks from 290 real-world repositories, offering a specialized alternative to general coding benchmarks. By focusing Antares on the initial triage stage of application security, Cisco is positioning small language models (SLMs) as efficient pre-filters for static and dynamic analysis tools, rather than complete replacements for the security stack. This follows a broader trend where agentic execution infrastructure funding rose to approximately $504 million in early 2026, signaling that investors and enterprises are moving beyond AI demos toward production-grade governance and monitoring systems.
Read full article at siliconangle.com
Get this in your inbox → Subscribe
Enjoy our coverage?
Add StreamingMeme as a preferred source on Google to see more of our streaming news at the top of your Search results.
Add as preferred source