Score
Designs, builds, and analyzes communication network architectures and their control mechanisms, including physical and logical topologies, switching/fabric designs, protocol stacks, software-defined control, and user-space or offload implementations. Develops and evaluates topology and flow models, placement and provisioning strategies, and optimization techniques for bandwidth, latency, routing, and host–network co-design to meet performance, scalability, and reliability requirements in high-speed and hybrid network deployments.
This work addresses the lack of systematic educational resources in high-performance computing (HPC) networking, which poses a significant barrier for researchers entering the field. It presents the first comprehensive integration of the HPC networking stack, covering communication protocol layers, programming interfaces such as MPI, control plane mechanisms, high-speed interconnect technologies, and custom link-layer hardware. The exposition is anchored by a detailed case study of the El Capitan supercomputer architecture at Lawrence Livermore National Laboratory. By offering a well-structured, practice-oriented primer, this contribution fills a critical educational gap and substantially lowers the entry barrier for researchers seeking to master core HPC networking technologies.
This work proposes a resource-aware efficiency metric for interconnection networks that explicitly accounts for the hardware overhead required to sustain non-blocking communication—specifically, link cost (α), crossbar cost (β), and concentration ratio—rather than focusing solely on latency or throughput. By modeling the impact of traffic hop count and router radix on these costs, the study systematically evaluates the cost optimality of various topologies under non-blocking constraints. It reveals that direct networks with high radix are superior for small to medium scales, whereas indirect topologies such as fat trees become necessary at large scales to manage router complexity. Furthermore, the analysis demonstrates that multi-plane star architectures achieve efficient fault tolerance with lower resource overhead compared to topologies relying on structural redundancy.
This study addresses the coordination challenges in in-band SDN control plane deployments with multiple controllers, where controller discovery, state synchronization, and failure recovery must be achieved without expanding switch forwarding state. The authors propose a boundary-switch-based local forwarding graph mechanism that confines inter-domain routing information to boundary devices, preventing state propagation into intermediate domains. In-band control communication is realized using Open vSwitch’s Nicira extensions with NSH encapsulation, and neighbor discovery is accomplished via Controller Advertisement messages. The approach requires no switch firmware modifications and incurs flow table overhead independent of the number of controllers, maintaining constant space complexity. Experiments in a Mininet environment with 96 switches and 5 controllers demonstrate that internal switches exhibit fixed flow table occupancy, enabling network scalability to hundreds of nodes with controller discovery convergence times on the order of seconds.
In enterprise network engineering, physical topology modifications and device configuration updates have long relied on error-prone, inefficient manual processes; existing automation research predominantly focuses on configuration synthesis while neglecting co-evolution with topology changes. This paper proposes the first intent-driven, closed-loop automation framework tailored for enterprise networks. It integrates multimodal large language models (MLLMs), optical character recognition (OCR), and a graph-structure-aware visual encoder to jointly understand topology diagrams and textual intent specifications. We introduce a novel topology–configuration co-prompting engineering paradigm and a Cisco-certified scenario fine-tuning mechanism. Evaluated on real-world enterprise deployments, our framework achieves significantly improved topology image parsing accuracy, reduces network design cycle time by over 40%, and attains an 89.2% execution accuracy for topology-modification intents—substantially decreasing manual intervention.
To address slow response to dynamic traffic demands and excessive topology reconfiguration oscillations in reconfigurable data centers, this paper proposes a batched dynamic graph scheduling algorithm based on incremental matching. It is the first work to introduce dynamic graph algorithms into optical circuit-switched network scheduling, modeling edge-disjoint matchings while explicitly capturing spatiotemporal locality of traffic to avoid full recomputation. We design six efficient batched algorithms and evaluate them on 176 synthetic and 39 real-world traffic traces. Compared to static approaches, our method reduces runtime by 92%, decreases configuration changes by 87%, and incurs zero loss in matching weight—achieving millisecond-scale responsiveness and high configuration stability. The core contribution lies in the principled integration of dynamic graph theory with optical switching characteristics, jointly optimizing update efficiency, reconfiguration stability, and matching quality.
This work addresses the core challenge in network automation: automatically generating deployable network topologies from natural language requirements while satisfying structural and resilience constraints. We propose a large language model (LLM)-based, constraint-driven framework that translates natural language into compliant topologies through hierarchical intent parsing and systematic validation. To facilitate evaluation, we introduce the first benchmark for this task, releasing a public dataset encompassing four real-world scenarios and characterizing common generation error patterns. Extensive experiments across multiple proprietary and open-source LLMs demonstrate the framework’s effectiveness, with performance quantified using metrics including topological correctness, node/edge F1 scores, and server-content connectivity. Our results provide actionable guidance for model selection in AI-driven network design.
This work addresses the lack of a high-throughput, general-purpose solution for large-scale AI training that simultaneously respects physical constraints and optimizes network topology, routing, and collective communication. The authors propose TONS, a framework that enables automated throughput-optimized network synthesis for AI supercomputers. TONS formulates topology synthesis as a linear optimization problem and scales to thousands of nodes by integrating theoretical insights with heuristic methods. It also introduces a deadlock-free routing mechanism supporting limited virtual channels and fault tolerance in optical switching. Under realistic deployment constraints, TONS achieves geometric mean speedups of 2.1× and 1.6× over the best TPU v4/5p torus variants for uniform random and all-to-all communication patterns, respectively.
This work addresses the challenges of in-band SDN control planes in resource-constrained wide-area telecommunication networks, including autonomous bootstrap, source routing, sub-50ms failure recovery, and multi-controller coordination. The authors propose Periplus, a system that embeds a forwarding graph—encoding both primary and per-hop backup paths—into L2/L3 packet headers. This design enables controller-independent local failover within 50 milliseconds and facilitates switch bootstrap with minimal flow table overhead. Notably, only two switches require initial configuration, eliminating network-wide flooding during provisioning. Flow table occupancy is decoupled from network size, scaling only at nodes that encode multipath information. Experimental evaluation using Ryu and Open vSwitch (augmented with Nicira extensions for NSH encapsulation) demonstrates Periplus’s capabilities in rapid recovery, scalable bootstrapping, and efficient resource utilization.
This study addresses the lack of efficient resource allocation strategies for HyperX networks, a challenge exacerbated by the inapplicability of existing methods from other topologies. The work presents the first systematic investigation of this problem, formally defining three classes of allocation strategies—linear, geometric, and random—and analyzing them theoretically through key topological properties such as expansion, convexity, and partition bandwidth. Extensive simulations are conducted under multiple routing algorithms using both synthetic traffic patterns and application-derived communication kernels. The results demonstrate that the non-convex Diagonal strategy consistently outperforms conventional approaches across most scenarios, revealing that partition bandwidth and switch locality are critical factors in mitigating interference and enhancing communication performance. These insights yield practical guidelines for deploying high-performance computing systems based on HyperX interconnects.