arXiv · 2609.37243
Codebook-Guided Cross-Modal Knowledge Distillation for Structurally Heterogeneous Features
Abstract
Cross-modal knowledge distillation transfers knowledge from a teacher modality to a student modality. Existing feature-level alignment methods typically assume that teacher and student features reside in structurally alignable representation spaces. However, this assumption does not hold when cross-modal features are structurally heterogeneous and lack clear unit-level correspondence, such as 2D spatial visual grids and 1D temporal audio sequences, thereby limiting the applicability of feature-level alignment. To address this challenge, we propose a cross-modal distillation framework that enables effective knowledge transfer across structurally heterogeneous feature spaces via a vector-quantized codebook. Specifically, teacher features are abstracted into a set of vector-form codes regardless of their original feature structure, and the selected codes serve as concept-level anchors for student learning. Code selection is guided by both task relevance and student compatibility, allowing the student to receive transferable teacher knowledge without requiring direct unit-level feature alignment. Experimental results across diverse cross-modal distillation scenarios demonstrate the effectiveness of the proposed framework on classification and semantic segmentation tasks.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Dae Ung Jo, Jongin Lim, YoungJoon Yoo, Daeho Um. 2026-09-29. Codebook-Guided Cross-Modal Knowledge Distillation for Structurally Heterogeneous Features. https://arxiv.org/abs/2609.37243
Cite the original work for its findings. Save a collection to share your selection of sources.