TrackFlood: Relocating Latency Attacks from NMS-Free Detectors to Real-Time Trackers
We consider latency attacks on object detectors, where the attacker's goal is not to corrupt a prediction but to make the system fail to respond in time, targeting real-time applications such as autonomous driving. Modern object detectors eliminate Non-Maximum Suppression (NMS) through one-to-one assignment or set prediction, removing the classical detector-side latency bottleneck exploited by prior latency (``sponge'') attacks. We show that NMS-free does not mean latency-robust: this architectural change does not eliminate the attack surface but relocates it downstream to multi-object tracking, whose data-association cost remains content dependent. We present \emph{TrackFlood}, a unified white-box overload attack against NMS-free detect-then-track pipelines spanning both one-to-one detectors (YOLOv10 and YOLO26) and query-based detectors (RT-DETR). TrackFlood recovers differentiable confidence tensors and optimizes perturbations that flood the tracker with spatially distributed phantom detections while leaving detector inference unchanged. Evaluated entirely on an NVIDIA Jetson AGX Orin (TensorRT FP16), detector latency remains essentially constant, whereas tracker latency increases substantially. At a standard imperceptible budget ($L_\infty{=}8/255$), a universal perturbation produces clearly measurable tracker overload but only modest end-to-end slowdown, without deadline misses for the association-dominated trackers; a separate higher-budget stress test drives severe end-to-end slowdowns and sustained deadline misses. We further evaluate a lightweight, architecture-agnostic bounded-admission layer that caps the tracker workload and largely restores end-to-end latency, at a non-trivial cost in admitted clean detections. Our results demonstrate that evaluating NMS-free perception systems requires considering downstream tracking and end-to-end timing, not detector inference alone.