arXiv Science⌕ Search

arXiv · 2610.06530

Improved LZ77 Compression with Match-Length-Dependent Sliding Windows

Abstract

We devise and analyze WLZ, a family of LZ77 encoders whose sliding-window sizes depend on match length: short matches use smaller windows and shorter distance fields, while long matches retain access to distant repetitions. Write $B:=\log_2 W$ for the maximum window $W$. A parsing-transfer bound charges both window restrictions and missed matches, preserving established convergence and finite-input minimax orders, and a sharper block charge proves stationary-ergodic universality of exact greedy WLZ when full-window matching begins at length $o(B)$. A calibrated window schedule never increases minimum token cost relative to a single-window baseline and saves $Ω(BW^{-β})$ in rate on specified iid sources, with $β>2$. The remaining gains come from the token code rather than the windows. On a uniform iid source, a growing phrase cap with a cheap length symbol improves redundancy from the $Ω(\log B/B)$ of the original 1977 fixed-field format, whatever phrase cap it uses, to $O(\log\log B/B)$. Recoding a single-symbol run $(r,1)$ as $(1,r)$, with a logarithmic-cost count $r$ that may exceed the match cap, codes inputs with $O(n^α)$ runs, $0\leα<1$, in $O(n^α\log n)$ bits, versus $Ω(n)$ for that format with phrase cap $Θ(\log W)$ and $W=o(n)$. On globally $p$-periodic inputs, Huffman coding the fields of a capped parse with nearest-distance ties reduces the large-file rate from $Θ(B/W)$ under fixed-width coding to at most $3/W$, including tables and framing. Finally, selection among complete WLZ codes yields an entropy-rate estimator consistent almost surely and in mean, with finite-data error bounds for fixed nonuniform iid sources.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Yingquan, Wu. 2026-10-05. Improved LZ77 Compression with Match-Length-Dependent Sliding Windows. https://arxiv.org/abs/2610.06530

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Parametric and structure-aware information theory for multi-scale analysis of composition

Compositional data is common across the natural and social sciences, requiring methods to measure the diversity of compositions and the heterogeneity of collections of them. Via Bregman geometry, a strictly concave diversity index induces a measure of heterogeneity that decomposes across scales. A canonical example is Shannon entropy inducing mutual information. We develop this framework by incorporating category similarity into $α$-logarithmic entropy, a parametric diversity index. We disprove an existing general concavity result and prove that, for $α=3$, the region of strict concavity is always convex, complementing previous results for $α=1$ and $2$. The resulting geometry provides methods to compare compositions substantially faster than optimal transport. Applied to occupation compositions across England and Wales, they reveal distinct regionalisations based on different notions of occupation relatedness. Applied to ecological data, they recover established patterns of functional and taxonomic $β$-diversity, and reveal sensitivity to the diversity parameter.

cs.IT↗

Improved Characterization of the Memory-Rate Tradeoff for Demand-Private Coded Caching With Multiple Demands

We consider a coded caching problem with multiple demands under a privacy constraint. In this problem, a server with access to $N$ files serves $K$ users over a shared link, and each user requests $L$ distinct files. The privacy constraint requires that each user obtain no information about the demands of the other users. We first propose a new achievable scheme for arbitrary $N$, $K$, and $L$ by applying a novel transformation to a suitable non-private coded caching scheme. We then derive a new converse bound, and show that the proposed scheme is order optimal within a multiplicative factor of $6$ of this bound. Finally, for the case of two users, we completely characterize the optimal memory-rate tradeoff for arbitrary $N$ and $L$ through three new achievable schemes and matching converse bounds.

cs.IT↗

Learning LDPC codes with density evolution over relaxed protographs

We consider the design of low-density parity-check (LDPC) codes for a given iterative decoder. While LDPC performance can be evaluated using simulation, density evolution (DE), or EXIT-chart analysis, selecting a parity-check matrix (PCM) remains a difficult combinatorial optimization problem. Existing approaches often rely on population-based search, random mutations, or genetic algorithms, which require careful tuning and incur high computational cost. Recent gradient descent (GD)-based methods optimize relaxed PCMs by differentiating through decoder simulations, but rely on noisy Monte Carlo estimates, line searches over soft matrix representations, and remain costly for long codes. Moreover, the loss is typically evaluated only at integer-valued PCMs. We focus on long protograph-based LDPC codes and propose a deterministic GD-based framework operating directly on a relaxed protograph representation. The loss is based on DE bit error rate (BER) and can be evaluated directly for relaxed protographs. To justify this relaxation, we associate the relaxed representation with an ensemble of binary PCMs and show that the proposed relaxed DE yields the ensemble-averaged DE performance. The resulting procedure supports standard GD optimization and achieves fast, reliable convergence through deterministic DE evaluation and informative gradients. Numerical results show that the optimized protographs outperform 5G LDPC codes with matching dimensions.

cs.IT↗