🤖 AI Summary
This study addresses the lack of efficient resource allocation strategies for HyperX networks, a challenge exacerbated by the inapplicability of existing methods from other topologies. The work presents the first systematic investigation of this problem, formally defining three classes of allocation strategies—linear, geometric, and random—and analyzing them theoretically through key topological properties such as expansion, convexity, and partition bandwidth. Extensive simulations are conducted under multiple routing algorithms using both synthetic traffic patterns and application-derived communication kernels. The results demonstrate that the non-convex Diagonal strategy consistently outperforms conventional approaches across most scenarios, revealing that partition bandwidth and switch locality are critical factors in mitigating interference and enhancing communication performance. These insights yield practical guidelines for deploying high-performance computing systems based on HyperX interconnects.
📝 Abstract
As high-performance computing systems scale in size and complexity, efficient resource management is essential to minimize communication overhead. The HyperX is a richly connected, low-diameter network that offers a scalable and cost-effective alternative to traditional topologies. However, resource allocation in HyperX remains underexplored, and strategies designed for networks like Torus, Fat-tree, or Dragonfly do not directly transfer. In this work, we propose and formalize several resource allocation strategies for HyperX networks, categorized into linear, geometric, and stochastic functions. We characterize these strategies theoretically by analyzing their topological properties, including dilation, convexity, and partition bandwidth.Furthermore, we conduct an exhaustive experimental evaluation using synthetic traffic and application communication kernels to assess the impact of these strategies on performance under different routing algorithms. Our results indicate that partition bandwidth and switch locality are decisive factors in mitigating interferences. Notably, the Diagonal allocation strategy, which is not convex, consistently outperforms traditional approaches in most scenarios. Finally, we provide a set of lessons learned to guide the implementation of resource allocation policies in HPC systems based on HyperX networks.