arXiv · 2610.12139
Toward Intelligent Networks via AI-Aware GPU-Native Packet Processing
Abstract
The integration of inline Artificial Intelligence (AI) models into critical network infrastructure is fundamentally bottlenecked by the high latency and synchronization overhead of CPU-mediated packet processing. While legacy GPU offload and recent CPU-bypass frameworks attempt to bridge this gap, they remain trapped in proprietary ecosystems or still rely on the host CPU and coarse-grained batching to coordinate stateful telemetry and AI pipeline execution. In this paper, we propose AGP, a novel framework that promotes the GPU from a passive accelerator to a primary data-path controller, answering what changes architecturally when the GPU autonomously owns the complete packet-to-inference pipeline. By enabling GPU-native packet processing, contention-free in-GPU stateful aggregation, and a persistent mega-kernel for continuous AI inference, AGP removes the CPU from the critical path. Our evaluation demonstrates that native in-GPU packet processing sustains line-rate throughput while being over 6.1x more power-efficient. Furthermore, our integrated Intrusion Detection System (IDS) stress test eliminates the legacy CPU-GPU synchronization tax, reducing end-to-end latency by 7.9x (up to 35x p99) and accelerating whole-system inference throughput by 9.3x using only ~2% of the GPU's thread capacity.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Seyed Mohammad Mehdi Mirnajafizadeh, Yiwen Hu, Rhongho Jang. 2026-10-08. Toward Intelligent Networks via AI-Aware GPU-Native Packet Processing. https://arxiv.org/abs/2610.12139
Cite the original work for its findings. Save a collection to share your selection of sources.