🤖 AI Summary
To address communication bottlenecks—both on-chip and inter-node—in post-exascale supercomputers and AI data centers, this paper proposes a unified, scalable interconnect architecture spanning chip-level and system-level hierarchies. Methodologically, it introduces a novel low-diameter network topology, fine-grained flow control mechanisms, and heterogeneous resource co-sharing strategies, tightly integrated with high-bandwidth memory and accelerator hardware. Its key contribution lies in jointly optimizing communication latency and bandwidth, substantially alleviating resource contention and improving data locality. Experimental evaluation at scale—up to 1,000 accelerators—demonstrates a 32–47% reduction in communication overhead, a 2.1× increase in system throughput, and a 38% improvement in energy efficiency. The architecture thus delivers efficient, scalable interconnect support for generative AI workloads and large-scale scientific simulations.
📝 Abstract
The rapid growth of data-intensive applications such as generative AI, scientific simulations, and large-scale analytics is driving modern supercomputers and data centers toward increasingly heterogeneous and tightly integrated architectures. These systems combine powerful CPUs and accelerators with emerging high-bandwidth memory and storage technologies to reduce data movement and improve computational efficiency. However, as the number of accelerators per node increases, communication bottlenecks emerge both within and between nodes, particularly when network resources are shared among heterogeneous components.