arXiv Science⌕ Search

arXiv · 2609.35310

Configuration-Induced Delivery Failures in NATS JetStream: Detection and Remediation

Abstract

NATS JetStream's at-least-once delivery guarantee is conditional: five common configuration mistakes silently violate it, causing duplicate message processing, data loss, or redelivery storms with no error logged anywhere. The standard Prometheus NATS exporter exposes only server-level throughput metrics and cannot detect any of these failures. We present nats-lens, a standalone monitor that reads from the JetStream management API and detects all five violation classes without requiring changes to monitored applications or client code. We formally characterize each class with a precise condition, prove that standard Prometheus NATS metrics are structurally incapable of detecting any of them, and implement five targeted detectors. In a controlled evaluation of 30 rounds per scenario on both single-node and 3-node JetStream clusters, nats-lens achieves 100% detection coverage across all five classes---versus 0% for the baseline---with zero false positives over 30 minutes of healthy operation. Detection latency ranges from 2,003 ms to 8,013 ms (within three poll cycles). We confirm language-agnostic detection empirically using consumers in Rust, Go, and Python. The tool is open source and exposes findings through four output channels: web dashboard, Prometheus metrics, REST API, and NATS health events.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Biplab Kumar Das. 2026-09-28. Configuration-Induced Delivery Failures in NATS JetStream: Detection and Remediation. https://arxiv.org/abs/2609.35310

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Cell-Free Massive MIMO Under Mobility: A Fairness-Differentiated Handover Scheme

While cell-free massive MIMO (CF-mMIMO) offers high and uniform network-wide throughput in static networks, its performance in mobile networks is not yet fully addressed. In this paper, we evaluate the throughput performance of urban mobile CF-mMIMO networks under a comprehensive throughput model and show that it suffers from large performance degradation due to the combined effect of channel aging and handover overheads. To restore the uniformly good performance of CF-mMIMO under mobility, we formulate a novel optimization problem to maximize the nett throughput that considers both channel aging and handover cost. We derive a near-optimal solution nearOpt for our transformed and relaxed optimization problem with Newton's method. We then design a heuristic handover algorithm, FairDiff, to differentiate prioritized and optional handovers using a policy threshold based on Jain's fairness index, in order to achieve uniform throughput over the network. Our extensive evaluation of the mobile throughput performance of our handover schemes in realistic urban mobile networks shows that, unlike the existing literature benchmarks that obtain very low throughput under mobility, our FairDiff scheme consistently achieves the near-optimal throughput comparable to nearOpt and highest network-wide throughput with the lowest computational complexity among all considered schemes. We thus for the first time propose a handover scheme that delivers the promise of uniformly good throughput for mobile CF-mMIMO, making it a feasible architecture for practical mobile networks.

cs.NI↗

Embodied AI in 6G Networks: From Intelligent Connectivity to Physical Intelligence

Embodied artificial intelligence (AI) couples perception and learned decision making to actions that change the physical world. This coupling distinguishes an embodied agent from a conventional connected controller: the agent maintains task state and uncertainty, reasons about the consequences of actions, and adapts from subsequent observations. Wireless networking becomes relevant when perception, inference, or coordination is distributed, but it should not replace local safety control. This article develops a tutorial perception--communication--action (PCA) architecture that exposes task state, action deadlines, uncertainty, agent intent, and safety envelopes to a 6G orchestration plane. It separates capabilities already addressed by 5G and 5G-Advanced from functions that motivate 6G, including task-state interfaces, semantic freshness, predictive digital twins, and safety-aware coordination across agents. A multi-robot simulation study is retained to illustrate joint sensing, communication, and computation control. The results show where network orchestration improves task utility and where local autonomy remains essential.

cs.NI↗

HOCCL: Offloading Collective Communication from GPU Cores to Accelerate Distributed Training

Large language model training involves massive computation on GPU streaming multiprocessors (SMs), the primary compute units of GPUs. Since SMs host specialized accelerators such as Tensor Cores, their efficient utilization is critical to training efficiency. Unfortunately, existing collective communication systems compete with computation for SMs, as they consume SMs for communication-related data movement and synchronization operations. We observe that communication can, in principle, be driven by DMA engines, thereby eliminating SM involvement in communication. Based on this insight, we propose HOCCL, a zero-SM collective communication framework consisting of three components: a stream manager, a point-to-point (P2P) executor, and a collective scheduler. The stream manager preserves operator-level temporal ordering with other GPU kernels. The P2P executor enables zero-SM point-to-point communication, while the collective scheduler orchestrates P2P transfers to maximize bandwidth. Experiments show that HOCCL preserves near-peak communication performance, achieving within 3% of the state of the art on average, while eliminating communication occupancy on nearly 10% of total GPU SMs. By freeing SM resources for computation, HOCCL improves end-to-end training throughput by up to 5%.

cs.NI↗