arXiv · 2609.37225
ResComEmb: Effective and Efficient Multimodal Embedding via Residual Homogeneity Compression
Abstract
Multimodal large language models (MLLMs) have shown strong potential for universal multimodal representation learning. However, existing methods either compress each input into a single vector, limiting fine-grained expressiveness, or retain long sequences of visual-token vectors, incurring substantial storage and interaction costs. To resolve this trade-off, we propose ResComEmb, a trainable framework for effective and efficient universal multi-vector multimodal embedding. ResComEmb first encodes each input at native dynamic resolution into ordered global, intermediate, and fine-grained views. After MLLM contextualization and embedding projection, a trainable Residual Homogeneity Compression (RHC) module reduces within-granularity redundancy and cross-granularity repetition under explicit visual token budgets. Then, ResComEmb introduces a length-adaptive Bidirectional Late-Interaction Matching mechanism for robust query-document scoring, which averages the strongest token-level matches in each direction and combines the two scores using a weight based on how many valid tokens each side has. Extensive experiments on MMEB, ViDoRe V1, and ViDoRe V2 show that ResComEmb produces higher-quality universal multimodal embeddings than VLM2Vec-V2, and outperforms ColQwen2.5 in visual document retrieval using only 37.5% of its full visual token budget, demonstrating a favorable effectiveness-efficiency trade-off.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Zijing Cai, Yuzhe Wang, Jingxian Zhu, Fengbin Zhu, Richang Hong. 2026-09-29. ResComEmb: Effective and Efficient Multimodal Embedding via Residual Homogeneity Compression. https://arxiv.org/abs/2609.37225
Cite the original work for its findings. Save a collection to share your selection of sources.