arXiv · 2605.20689
DIVE: Embedding Compression via Self-Limiting Gradient Updates
Abstract
High-dimensional language-model embeddings increase storage and search costs, while supervised compressors can overfit when relevance labels are scarce. We present DIVE (Dimensionality reduction with Implicit View Ensembles), a residual compression adapter codesigned with a self-limiting hinge loss, geometry distillation, and head-wise NT-Xent over implicit coordinate views. The hinge stops updating satisfied ranking constraints, while the dense objectives stabilize the compressed representation; only the first head is retained at inference. Under query-disjoint evaluation with two LLM2Vec backbones, five BEIR benchmarks, 128d and 256d outputs, and six baselines, DIVE is the strongest adapter on all five primary benchmarks. It also outperforms PCA and an autoencoder in comparisons against unsupervised compressors.
Explore related subjects
Keep this discovery
Dongfang Zhao. 2026-05-20. DIVE: Embedding Compression via Self-Limiting Gradient Updates. https://arxiv.org/abs/2605.20689
Cite the original work for its findings. Save a collection to share your selection of sources.