Huawei представляет новую вычислительную архитектуру UnifiedBus для SuperPoD и кластеров
PRNewsWire
公開日時: Sep 22, 2026, 12:49 PM GMT+9
Sentiment Analysis
In traditional computing architecture, resource utilization decreases as cluster scale increases. Models actually use only 20% of the computing power of a 100,000-NPU cluster, with colossal computing resources lying idle during data transfer. A 10-trillion-parameter model, which requires massive amounts of intermediate data for training, far exceeds the memory capacity on a single accelerator in traditional architectures. Interconnection and communication between components in a cluster have become a key performance bottleneck for computing systems.
According to Yang, Huawei's solution to this bottleneck is a new computing architecture based on an interconnect technology called UnifiedBus, which facilitates interaction between clusters and SuperPoDs. This architecture was designed to support the computing demand boom that is growing as new agent applications emerge in the market.
Computing systems based on UnifiedBus have four key features:
Unified protocol and memory semantics: More than ten interconnect protocols have been merged into the UnifiedBus protocol, which increases interconnect bandwidth from 100 GB/s to TB/s and reduces round-trip time (RTT) from 7 microseconds to 2 microseconds. This protocol also enables unified global memory addressing in SuperPoD.
Heterogeneous computing collaboration: UnifiedBus directly connects CPUs, NPUs, memory, and solid-state drives (SSDs) to enable decentralized peer-to-peer access between them and a flexible CPU-NPU mix. Additionally, multi-level hardware acceleration for Transformer enables disaggregation of Attention and FFN (AFD) layers.
Multi-level storage with a global pool: UnifiedBus uses hybrid environment resource pooling to support activation caching. Double data rate (DDR) memory serves as alternative memory for NPUs, reducing lookup latency for recommendation and advertising services and doubling vector search performance for 100 billion data items with thousands of dimensions. This also reduces the HBM bandwidth requirements for each NPU for training 10-trillion-parameter models and improves the floating-point operations per second (FLOPS) utilization of the cluster model (MFU).
Optoelectronic communication and flexible networking: UnifiedBus serves as a global data highway with ultra-high bandwidth and ultra-low latency, enabling flexible computing scaling. Yang then introduced Huawei's latest UnifiedBus products, which cover intra-cabinet, inter-cabinet, and inter-cluster interconnects. This enables elastic scaling from a single cabinet to a cluster of one million NPUs.
Within cabinets, UnifiedBus LinkBlade eliminates link loss with a cable-free design, reducing copper cabling within a 4096-NPU SuperPoD by approximately 196 kilometers. Between cabinets, UnifiedBus LinkDevice supports 176 ports with 1.6 Tbps bandwidth per port, enabling a 280 Tbps all-optical connection. This high-speed bus protocol interconnect device has the highest bandwidth and the most ports in the industry, with an RTT of just 2 microseconds.
Across clusters, the UnifiedBus UBG switch...
Source: PRNewsWire
個別の投資に関する推奨やアドバイスを提供することを意図しておりません。ここで述べられている意見や見解は、あくまでも各記事の個人的見解です。