Score
Designs, implements, and analyzes the packet forwarding plane of network devices, including forwarding tables, lookup and fast‑path engines, packet processing pipelines, and interfaces to control plane and underlying hardware (ASIC/FPGA/CPU). Evaluates and optimizes forwarding correctness, policy enforcement, failure handling, and performance characteristics such as throughput, latency, and resource usage (memory, TCAM, pipeline stages).
This paper addresses the fundamental tension between hardware resource constraints—limited memory and restricted instruction sets—and functional complexity in programmable data planes (P4 switches) for network security. To resolve this, we propose an efficient implementation paradigm tailored for security functions. Our approach integrates loop-based processing, lookup-table precomputation, recirculate-and-truncate mechanisms, and pipeline-level optimizations to enable line-rate execution of DDoS and spoofing attack detection/mitigation, fine-grained next-generation firewall policy enforcement, in-network encryption, and lightweight machine learning inference. Contributions include: (1) a systematic characterization of the P4 security design space and the first identification of general-purpose optimization patterns for resource-constrained environments; (2) empirical validation of feasibility and performance bounds for multiple critical security capabilities on commercial P4 platforms; and (3) identification of promising future directions—including architecture co-design, security-semantic abstraction, and heterogeneous offloading.
Existing P4 network simulation platforms struggle to simultaneously achieve high fidelity and high performance, hindering the design and customization of programmable data-plane algorithms. This paper introduces the first ns-3-native, high-performance P4 simulation framework, featuring a novel deep integration architecture between the Behavioral Model v2 (bmv2) and ns-3. It supports multiple P4 target architectures—including V1Model, PSA, and PNA—and enables microsecond-level time synchronization, fine-grained queue modeling (e.g., PIFO and ECN-aware queues), and P4Runtime-driven host-switch co-simulation. Evaluated on Basic Tunneling and Load Balancing benchmarks, the framework achieves an equivalent throughput exceeding 100 Gbps—over one order of magnitude higher than prior solutions—while maintaining queue delay error below 1.2%. These advances significantly enhance both simulation fidelity and execution efficiency for programmable data-plane research and development.
To address performance bottlenecks—such as low packet-processing throughput and suboptimal link utilization—in data-plane programming (particularly P4) for SDN, this paper proposes a four-dimensional co-optimization framework: (1) extending the P4 language to support asynchronous external function calls; (2) designing a lightweight, load-size-adaptive compression mechanism; (3) introducing in-network caching to reduce redundant traffic overhead; and (4) offloading high-overhead functions to multi-threaded hosts, VMs, or remote servers. Experimental evaluation demonstrates that the approach preserves full P4 programmability while significantly improving throughput (average +37%), reducing end-to-end latency (up to −42%), and enhancing utilization of both network links and computational resources. This work establishes a systematic, deployable optimization framework for high-performance, highly flexible data-plane programming.
This work addresses the challenges posed by the lengthy and platform-specific nature of FPGA hardware design, which hinders flexible deployment in high-throughput, agile network infrastructures. To overcome this limitation, the authors propose PAF, a framework that employs a pipeline-oriented, parameterized architectural design methodology using the Chisel hardware construction language. By decoupling architectural intent from low-level implementation details, PAF enables fine-grained hardware control while substantially reducing code complexity. The framework facilitates efficient reuse and automated optimization of the same network application across diverse FPGA platforms. Experimental evaluation on an industrial-grade packet classification system demonstrates that PAF achieves performance and resource utilization comparable to hand-optimized designs, while significantly improving development productivity and portability.
Conventional network telemetry frameworks struggle to support fine-grained traffic measurement, performance diagnostics, and attack detection under stringent memory and computational constraints of high-speed network devices. Method: This paper proposes a lightweight, real-time online telemetry framework that systematically integrates compact data structures—including Bloom filter variants, Count-Min Sketch, and HyperLogLog—with streaming algorithms, hierarchical sampling, and P4-programmable data-plane co-design to comply with hardware limitations. Contribution/Results: Evaluated at line rate exceeding 100 Gbps, the framework reduces memory footprint by over 60% compared to state-of-the-art approaches while maintaining sub-1% flow frequency estimation error. It achieves an optimal trade-off among accuracy, throughput, and resource overhead, thereby significantly enhancing the feasibility and practicality of telemetry in high-bandwidth environments.
本文针对FPGA中NoC拥塞问题,通过集成拥塞成本、转弯模型路由算法及SAT求解方法,有效减少了95.1%的网络拥塞。
This work proposes Fastroute, a novel architecture addressing interconnect bandwidth bottlenecks and single-chip area limitations in multi-ASIC switches. Fastroute introduces an indirect layer that uniquely integrates packet and circuit switching, leveraging dynamic port remapping to localize traffic and substantially reduce cross-ASIC communication demands. A hardware prototype is developed and systematically evaluated using large language model training workloads. Experimental results demonstrate that Fastroute closely approximates single-ASIC performance while significantly reducing bandwidth overhead and power consumption. Furthermore, it effectively liberates external interface capacity to satisfy high-radix networking requirements, enabling system scalability without reliance on next-generation ASICs.
This study addresses the challenge that exact priority scheduling in programmable data planes struggles to achieve line-rate processing, while approximate approaches suffer from correctness deficiencies. To overcome these limitations, this work proposes two parallel architectures, P3PO and PPQA. The core innovation lies in decomposing a global priority queue into multiple parallel smaller queues, introducing the first parallel priority queue design based on cascaded boundaries or occupancy distributions to transcend the performance bottleneck of conventional single large-capacity queues. Leveraging the NetBench simulator, the project implements dynamic multi-queue multiplexing and evaluates it under web search and data mining workloads. Results demonstrate that the proposed approach effectively balances scheduling accuracy with hardware scalability, significantly enhancing scheduling precision in high-throughput scenarios.
This study addresses the challenge of dynamically coordinating heterogeneous traffic demands by a single agent in multi-engine analysis platforms. We propose an engine-agnostic traffic orchestration architecture based on a Composite Finite State Machine (CFSM). By modeling the multi-engine state space via Cartesian product and pre-mapping forwarding configurations, the method enables runtime table-lookup scheduling. Furthermore, temporal expiration is introduced as a first-class FSM transition to support declarative traffic decay, while a per-state hierarchical output mechanism is designed to accommodate varying service levels. The proposed architecture scales seamlessly to an arbitrary number of engines without structural refactoring, significantly enhancing both the efficiency and flexibility of traffic orchestration.
This study addresses cascading failures caused by shared firmware in commercial multi-host network interface cards (NICs) and the operational challenges of hyperscale deployments by proposing fbnic, a system-level solution. Architecturally, it introduces physical isolation and driver-priority mechanisms, combined with sub-sled-granularity firmware upgrade orchestration and slice-level fault containment. Operationally, it establishes a hardware-in-the-loop (HIL) continuous integration pipeline alongside cross-layer fault attribution monitoring and an automated remediation toolchain. Deployed across hundreds of thousands of hosts, fbnic reduces unplanned unavailability by 12×, shortens mean time to repair by 37%, and decreases hardware replacement rates by 2.3×, demonstrating robust stability at scale.