arXiv · 2609.35310
Configuration-Induced Delivery Failures in NATS JetStream: Detection and Remediation
Abstract
NATS JetStream's at-least-once delivery guarantee is conditional: five common configuration mistakes silently violate it, causing duplicate message processing, data loss, or redelivery storms with no error logged anywhere. The standard Prometheus NATS exporter exposes only server-level throughput metrics and cannot detect any of these failures. We present nats-lens, a standalone monitor that reads from the JetStream management API and detects all five violation classes without requiring changes to monitored applications or client code. We formally characterize each class with a precise condition, prove that standard Prometheus NATS metrics are structurally incapable of detecting any of them, and implement five targeted detectors. In a controlled evaluation of 30 rounds per scenario on both single-node and 3-node JetStream clusters, nats-lens achieves 100% detection coverage across all five classes---versus 0% for the baseline---with zero false positives over 30 minutes of healthy operation. Detection latency ranges from 2,003 ms to 8,013 ms (within three poll cycles). We confirm language-agnostic detection empirically using consumers in Rust, Go, and Python. The tool is open source and exposes findings through four output channels: web dashboard, Prometheus metrics, REST API, and NATS health events.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Biplab Kumar Das. 2026-09-28. Configuration-Induced Delivery Failures in NATS JetStream: Detection and Remediation. https://arxiv.org/abs/2609.35310
Cite the original work for its findings. Save a collection to share your selection of sources.