arXiv · 2609.01613
Skim and Skip: Hierarchical Adaptive Inference for Efficient Multimodal Retrieval
Abstract
Universal multimodal retrieval (UMR) increasingly adopts multimodal large language models (MLLMs) as unified embedding backbones, but their strong retrieval performance comes at substantial inference cost. Existing methods typically rely on uniformly dense inference, where all input tokens are processed through the entire model and matched using the final-layer [EOS] representation. However, this paradigm overlooks two key forms of heterogeneity in multimodal retrieval: token contributions to the final retrieval embedding are highly uneven, and different queries require markedly different amounts of inference depth. To address this, we propose Skim and Skip (SAS), a hierarchical adaptive inference framework for efficient multimodal retrieval. SAS first performs token-level evidence selection to preserve only the input information most relevant to the final retrieval embedding, and then performs depth-adaptive inference to determine whether the current representation is already sufficient for reliable matching. Experiments on 12 MMEB retrieval tasks show that SAS retains about 99% of the dense baseline's average retrieval performance while achieving up to 1.64 times end-to-end speedup and up to 66.3% FLOPs reduction.
Explore related subjects
Keep this discovery
Meng Gao, Yizhen Zhang, Yang Ding, Ziqi Dai, Shuoshuo Zhang, Junjie Wang, Taiqiang Wu, Chufan Shi, Lei Ji, Jian Jiao, Linfeng Zhang, Yeyun Gong, Yujiu Yang. 2026-06-24. Skim and Skip: Hierarchical Adaptive Inference for Efficient Multimodal Retrieval. https://arxiv.org/abs/2609.01613
Cite the original work for its findings. Save a collection to share your selection of sources.
Discover connections
Connections use source metadata and explicit phrase matches, not verified experimental comparisons.