LatencyLab: A DPDK-Based P4 Pipeline Latency Measurement Framework for FPGA SmartNICs

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of precisely measuring P4 pipeline latency in FPGA-based SmartNICs by proposing a DPDK-based, clock-synchronization-free measurement framework. The method leverages dual-port bypass comparison and CPU TSC timestamps to isolate pipeline latency without requiring PHC or PTP support, achieving hardware-level, low-overhead precision through network-segment multicast. Implemented on an AMD Alveo U280 platform using VitisNetP4, the system demonstrates that latency distributions are tight and reproducible, with 99% of packets experiencing delays under 20 ns. These results align closely with both independent verification methods and vendor-specified bounds.
📝 Abstract
P4-programmable FPGA SmartNICs place packet processing directly on the wire, but open FPGA P4 toolflows do not expose timestamping at the pipeline boundary, so the latency a P4 program adds on the target FPGA is rarely measured. This paper presents LatencyLab, a DPDK-based measurement framework for FPGA P4 pipeline latency that needs neither PHC/PTP support on the datapath nor clock synchronization. The FPGA's two ports share a network segment, so the switch multicasts a copy of each probe packet to both: one copy passes through the VitisNetP4 pipeline, the other through a matched bypass path. A kernel-bypass DPDK receiver busy-polls both ports and timestamps every packet with the CPU timestamp counter (TSC) as it is retrieved from the NIC's receive circular buffer. The arrival-time difference of the two copies isolates the pipeline latency after calibration against a null bitstream carrying the same traffic; transmit time cancel in the subtraction. We evaluate four VitisNetP4 programs on an AMD Alveo U280, probing each with a 20,000-packet trace measured ten times per session over five independent sessions, all TSC-timestamped and reflected for hardware timestamping. The measured latency distributions are tight and reproducible: 99% of packets fall within 20 ns of the median, session medians repeating within 1 to 2 ns (FiveTuple 107/137 ns, Forward 149 ns, RemoveHeader 177 ns, Checksum 364 ns at 250 MHz). Two independent checks agree with the framework: a kernel-free reflector returns every probe pair to a ConnectX-5 NIC whose adapter clock reproduces the measured distributions within a few nanoseconds, quantile by quantile, and every measured packet falls 18 to 22 clock cycles below the vendor's worst-case latency bound.
Problem

Research questions and friction points this paper is trying to address.

FPGA SmartNICs
P4 pipeline latency
latency measurement
timestamping
Innovation

Methods, ideas, or system contributions that make the work stand out.

DPDK
P4 pipeline latency
FPGA SmartNICs
TSC timestamping
VitisNetP4
🔎 Similar Papers
No similar papers found.