Huawei and Cisco lead IETF Fast Network Notifications for AI clusters
A group of industry engineers has submitted an IETF draft proposing a Fast Network Notification (FANN) framework to address sub-millisecond latency requirements in AI training and cloud clusters. The proposal aims to bridge the gap between hardware-level fault detection and control-plane responses to improve network reliability and performance.
Key Takeaways
- Proposed framework targets notification delivery in the order of milliseconds or sub-milliseconds to prevent transient network overload.
- Engineers from Huawei, Cisco, Tencent, and HPE identified that existing BFD and FRR mechanisms often take tens of milliseconds to detect failures.
- The draft addresses specific risks to AI training, where fiber link failures can cause hours of compute waste and energy loss.
- Proposed notifications would carry granular data including event type, queue buildup, and specific flow identification.
Why It Matters
The proposal addresses a critical bottleneck in distributed AI training where traditional routing protocols are too slow to prevent GPU synchronization barriers from being missed. By enabling sub-millisecond alerts, the framework allows the network layer to shift traffic before congestion or micro-loops degrade model convergence. For the streaming ecosystem, this technical development signals a shift toward more responsive data center interconnects capable of handling the bursty, high-bandwidth demands of generative AI and cloud-based rendering. As these clusters scale, the ability to bypass slow control-plane distribution will be essential for maintaining high utilization rates. Watch for the IETF Working Group to define specific information models and delivery modes for these lightweight signaling methods.
Additional Context
The FANN proposal arrives amid a broader IETF effort to modernize network signaling for AI workloads. Huawei has been particularly active in this space, with the company's data communication division publishing multiple IETF drafts on AI fabric networking in 2025 and 2026 that address lossless Ethernet requirements for GPU clusters. Cisco, which co-authored the FANN problem statement alongside Huawei engineers, has separately been advancing its own Ultra Ethernet Consortium contributions, where the consortium released its 1.0 specification in June 2025 targeting AI and HPC network fabrics with goals of reducing tail latency and improving congestion control beyond what standard Ethernet provides. The overlap between FANN's sub-millisecond notification goals and Ultra Ethernet's transport-layer improvements suggests these efforts may converge or compete as standards mature.
On the business and deployment side, the FANN framework reflects growing operator interest in AI-ready network infrastructure. China Telecom, China Mobile, and China Unicom, all listed as contributing organizations to the draft, have been investing heavily in intelligent computing centers. China Telecom announced in early 2025 that it would expand its AI computing capacity to over 10 EFLOPS by year-end, while Telefonica and Turkcell, also named in the draft, represent the European and Middle Eastern operator interest in standardizing low-latency signaling for their own AI cluster deployments. Equinix, the only colocation provider among the contributors, has been expanding its interconnection platform to support AI workload adjacency across its global footprint, positioning its facilities as neutral meeting points where FANN-style notifications could reduce cross-tenant latency penalties.
From a technical standpoint, FANN's sub-millisecond target sits at the intersection of several active research areas. The Ultra Ethernet Consortium's congestion control mechanisms aim for similar latency reductions but operate at the transport layer, whereas FANN focuses on fault notification propagation at the forwarding plane. A 2025 study from researchers at Meta and Broadcom measured that GPU training jobs lose up to 30% of effective throughput when network faults take longer than 100 microseconds to propagate, providing quantitative motivation for the FANN approach. Huawei's own CloudEngine switches already implement proprietary fast-failover mechanisms that achieve sub-millisecond switchover, and the company demonstrated these capabilities at its HC 2025 conference in September 2025 as part of its Intelligent IP Network portfolio for AI data centers. The IETF standardization effort would make such capabilities interoperable across vendors, which is the core value proposition of the FANN working group.
Read full article at datatracker.ietf.org
Enjoy our coverage?
Add StreamingMeme as a preferred source on Google to see more of our streaming news at the top of your Search results.
Add as preferred source