🤖 AI Summary
This study addresses the high-latency bottleneck caused by CPU-mediated packet processing and the inability of conventional GPU offloading frameworks to jointly orchestrate state telemetry and AI inference. We propose AGP, a framework that restructures the GPU from a passive accelerator into an active data-path controller. Through GPU-native packet processing, on-chip lock-free state aggregation, and a persistent large-kernel architecture, AGP entirely eliminates CPU dependencies on the critical path, achieving fully autonomous control from packet reception to AI inference. Experimental results demonstrate that AGP maintains line-rate throughput while improving energy efficiency by over 6.1×. In intrusion detection system (IDS) scenarios, it reduces end-to-end latency by 7.9× (up to 35× at p99) and accelerates inference throughput by 9.3×, all while consuming merely approximately 2% of GPU thread resources.
📝 Abstract
The integration of inline Artificial Intelligence (AI) models into critical network infrastructure is fundamentally bottlenecked by the high latency and synchronization overhead of CPU-mediated packet processing. While legacy GPU offload and recent CPU-bypass frameworks attempt to bridge this gap, they remain trapped in proprietary ecosystems or still rely on the host CPU and coarse-grained batching to coordinate stateful telemetry and AI pipeline execution. In this paper, we propose AGP, a novel framework that promotes the GPU from a passive accelerator to a primary data-path controller, answering what changes architecturally when the GPU autonomously owns the complete packet-to-inference pipeline. By enabling GPU-native packet processing, contention-free in-GPU stateful aggregation, and a persistent mega-kernel for continuous AI inference, AGP removes the CPU from the critical path. Our evaluation demonstrates that native in-GPU packet processing sustains line-rate throughput while being over 6.1x more power-efficient. Furthermore, our integrated Intrusion Detection System (IDS) stress test eliminates the legacy CPU-GPU synchronization tax, reducing end-to-end latency by 7.9x (up to 35x p99) and accelerating whole-system inference throughput by 9.3x using only ~2% of the GPU's thread capacity.