arXiv Science⌕ Search

arXiv · 2610.12139

Toward Intelligent Networks via AI-Aware GPU-Native Packet Processing

Abstract

The integration of inline Artificial Intelligence (AI) models into critical network infrastructure is fundamentally bottlenecked by the high latency and synchronization overhead of CPU-mediated packet processing. While legacy GPU offload and recent CPU-bypass frameworks attempt to bridge this gap, they remain trapped in proprietary ecosystems or still rely on the host CPU and coarse-grained batching to coordinate stateful telemetry and AI pipeline execution. In this paper, we propose AGP, a novel framework that promotes the GPU from a passive accelerator to a primary data-path controller, answering what changes architecturally when the GPU autonomously owns the complete packet-to-inference pipeline. By enabling GPU-native packet processing, contention-free in-GPU stateful aggregation, and a persistent mega-kernel for continuous AI inference, AGP removes the CPU from the critical path. Our evaluation demonstrates that native in-GPU packet processing sustains line-rate throughput while being over 6.1x more power-efficient. Furthermore, our integrated Intrusion Detection System (IDS) stress test eliminates the legacy CPU-GPU synchronization tax, reducing end-to-end latency by 7.9x (up to 35x p99) and accelerating whole-system inference throughput by 9.3x using only ~2% of the GPU's thread capacity.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Seyed Mohammad Mehdi Mirnajafizadeh, Yiwen Hu, Rhongho Jang. 2026-10-08. Toward Intelligent Networks via AI-Aware GPU-Native Packet Processing. https://arxiv.org/abs/2610.12139

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

IPEK: Event-Criticality-Aware Evidential Trust Management against Strategic Insider Attacks in Vehicular Ad Hoc Networks

Event-based trust management in vehicular ad hoc networks is vulnerable to strategic insider attackers who report trivial events correctly to build reputation, falsify reports on safety-critical events, and corrupt the trust feedback they provide about other vehicles. This paper proposes IPEK, an event-criticality-aware evidential trust management framework. IPEK combines an asymmetric local trust that scales rewards and penalties with event severity and location criticality, a reporter-credibility test that discounts feedback deviating systematically from the consensus of witnesses, and a Dempster-Shafer global trust that fuses feedback with Yager's rule, historical combination and risk accentuation. Closed-form properties quantify how quickly trust is lost and regained and show why Yager's rule keeps the global state revisable. IPEK is evaluated in simulation against a trust cascading-based dissemination scheme and the beta reputation system under eight attacks, with severity-, location-, on-off- and continuously triggered attackers falsifying their reports alone or also their feedback, at attacker ratios from 10% to 40%. Threshold-free separation measures are complemented by closed-loop revocation with frozen thresholds. With up to 25% attackers, IPEK achieves a significantly higher F1-score than both baselines under all attacks, with false positive rates of 0.9-15.5% against 7.8-49.7% and 26.0-41.4%, and it remains robust up to 40% attackers when only reports are falsified. Under fully colluding attackers that also corrupt feedback, all methods break down beyond about one third of attackers; this envelope and a longer detection delay are reported as limitations.

cs.NI↗

SparKV: Overhead-Aware KV Cache Loading for Efficient On-Device LLM Inference

Efficient inference for on-device Large Language Models (LLMs) remains challenging due to limited hardware resources and the high cost of the prefill stage, which processes the full input context to construct Key-Value (KV) caches. We present SparKV, an adaptive KV loading framework that combines cloud-based KV streaming with on-device computation. SparKV models the cost of individual KV chunks and decides whether each chunk should be streamed or computed locally, while overlapping the two execution paths to reduce latency. To handle fluctuations in wireless connectivity and edge resource availability, SparKV further refines offline-generated schedules at runtime to rebalance communication and computation costs. Experiments across diverse datasets, LLMs, and edge devices show that SparKV reduces Time-to-First-Token by 1.3$x-5.1x with negligible impact on response quality, while lowering per-request energy consumption by 1.5x to 3.3x, demonstrating its robustness and practicality for real-world on-device deployment.

cs.NI↗

Hardware-Software Co-Design for Energy-Efficient Neural Recording in Wireless Implantable BCIs

Implantable Brain-Computer Interfaces (iBCIs) are increasingly pivotal in clinical and daily applications. However, wireless iBCIs face severe constraints in power consumption and data throughput. To mitigate these bottlenecks, we propose a wireless iBCI headstage featuring adaptive ADC sampling and spike detection. Distinguishing our design from traditional application-layer compression, we employ a server-driven architecture that achieves source-level efficiency. Specifically, the server learns an optimal, electrode-specific sample rate vector to dynamically reconfigure the ADC hardware. This strategy reduces data volume directly at the acquisition layer (ADC and amplifier) rather than relying on computationally intensive post-digitization processing. Extensive experiments across diverse subjects and arrays demonstrate a power reduction of up to 40 mW and a 3.2x decrease in FPGA resource utilization, all while maintaining or exceeding decoding accuracy in both motor and visual tasks. This design offers a highly practical solution for long-term in-vivo recording.

cs.NI↗