arXiv Science⌕ Search

arXiv · 2609.32242

Curv-Tail: Lightweight Long-Tailed Encrypted Traffic Classification with Discrete Packet-Length Encoding and Lorentz Prototypes

Abstract

Long-tailed encrypted traffic classification requires accurate recognition of infrequent classes under limited computational budgets. We propose Curv-Tail, a lightweight, end-to-end packet--byte framework trained without a separate pretraining stage. Mixed-resolution tokenization preserves exact packet-length identities within a bounded range and coarsens larger values to limit the vocabulary. An auxiliary objective predicts observed length tokens from contextual packet features before pooling, encouraging length-token retention beyond flow-level supervision. Compact temporal encoders process packet sequences and directional byte patches, and Lorentz prototypes with a shared learnable curvature magnitude classify their fused representation. On NUDT-Mobile and DataCon-Website under natural class frequencies, Curv-Tail achieves three-seed mean Tail-F1 scores of 85.70% and 46.30%, exceeding the strongest evaluated baselines by 2.91 and 1.89 percentage points, respectively. In 300-class profiling on an RTX 4090 with FP32 and batch size 256, Curv-Tail uses 98.48% fewer parameters and achieves 10.2 times the batch inference throughput of MM4Flow.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Yankun Wang, Jun-Jie Huang, Lin Liu, Xiaodong Lei, Yi Chen, Yeqing Yan, Jiangyong Shi, Yongjun Wang. 2026-09-26. Curv-Tail: Lightweight Long-Tailed Encrypted Traffic Classification with Discrete Packet-Length Encoding and Lorentz Prototypes. https://arxiv.org/abs/2609.32242

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Scalable Long-Term Beamforming for Massive Multi-User MIMO

Fully digital massive multiple-input multiple-output (MIMO) systems with large numbers (1000+) of antennas offer capacity gains from spatial multiplexing and beamforming, but receivers that scale to these array dimensions face challenges in both channel estimation overhead and digital computation. Long-term beamforming addresses both by projecting the data onto a low-dimensional subspace that can be tracked at a slow time scale from the long-term channel parameters. In this setting, we show how to compute, in closed form, the projection matrix that maximizes a capacity upper bound, using a matrix inverse square root; the same projection is shown to maximize the mean post-projection signal-to-interference-plus-noise ratio (SINR) exactly. Computationally efficient methods are then presented for the matrix computation, realizable with matrix-matrix multiplies and hence amenable to systolic array implementations in hardware. Bounds on the SINR degradation are derived, and ray tracing simulations in a realistic rural uplink setting show a small loss relative to instantaneous minimum mean-square error (MMSE) beamforming when the covariance is accurately estimated. The efficient Gram-domain form of the instantaneous MMSE receiver applies the maximum-ratio combining reduction before an inverse whose dimension is the total number of streams. Against this baseline, the method estimates and refreshes beamforming coefficients three orders of magnitude less often and decouples the real-time path across users. With the conjugate-gradient solve and a rank-one projection, its total arithmetic cost is 2% higher at the ten-user operating point and lower above a stream-dimension crossover that we characterize.

eess.SP↗

Airborne Particle Communication Through Time-varying Diffusion-Advection Channels

Particle-based communication using diffusion and advection has emerged as an alternative signaling paradigm recently. While most existing studies assume constant flow conditions, real macro-scale environments such as atmospheric winds exhibit time-varying behavior. In this work, airborne particle communication under time-varying advection is modeled as a linear time-varying (LTV) channel, and a closed-form, time-dependent channel impulse response is derived using the method of moving frames. Based on this formulation, the channel is characterized through its power delay profile, leading to the definition of channel dispersion time as a physically meaningful measure of channel memory and a guideline for symbol duration selection. System-level simulations under directed, time-varying wind conditions show that waveform design is critical for performance, enabling multi-symbol modulation using a single particle type when dispersion is sufficiently controlled. To quantify waveform distortion and guide the design of orthogonal signaling waveforms, the Orthogonality Loss Ratio (OLR) is introduced as a structural metric. The results demonstrate that time-varying diffusion-advection channels can be systematically modeled and engineered using communication-theoretic tools, providing a realistic foundation for particle-based communication in complex flow environments.

eess.SP↗

Extended Universal Joint Source-Channel Coding for Digital Semantic Communications: Improving Channel-Adaptability

Deep learning (DL)-based joint source--channel coding (JSCC), particularly vector quantization (VQ)-based JSCC, enables efficient semantic communication by mapping high-dimensional feature vectors into compact codeword indices for digital modulation. However, existing methods, including universal JSCC (uJSCC), rely on fixed, modulation-specific encoders, decoders, and codebooks, limiting adaptation to fine-grained channel variations. We propose an extended universal JSCC (euJSCC) framework that enables adaptive transmission across both modulation orders and fine-grained channel conditions within a single model using hypernetwork-based normalization and a dynamic codebook generation (DCG) network that refines modulation-specific base codebooks according to channel quality indicator (CQI) feedback. To support time-varying channels with periodic CQI feedback, euJSCC employs an inner--outer encoder--decoder architecture, where the outer modules capture long-term channel statistics and the inner modules refine feature vectors to align with block-wise codebooks. A two-phase training strategy first pretrains the outer modules and DCG network over additive white Gaussian noise (AWGN) channels to establish stable feature--codebook alignment and then fine-tunes the full model over time-varying channels. Experiments demonstrate that euJSCC consistently outperforms state-of-the-art channel-adaptive digital JSCC schemes in image transmission, while additional evaluations under practical channel and feedback conditions and on a downstream task further validate its effectiveness.

eess.SP↗