arXiv · 2609.23194
Enhancing speech representation learning with cross-modal knowledge transfer with HGNN under low resource settings: the case study of Yemba
Abstract
Acoustic representation learning is crucial for speech processing, yet low-resource languages (LRLs) face severe data scarcity, limiting the effectiveness of traditional and self-supervised methods. As a promising alternative, in this work, we propose to enhance acoustic representation trough a cross-modal transfer knowledge approach, based on heterogeneous graph neural networks (HGNNs), where acoustic and linguistic entities are modeled as distinct node types within a unified graph. Through message-passing mechanisms, linguistic nodes explicitly transfer knowledge to acoustic nodes, enabling structured and interpretable cross-modal information flow. To highlight this knowledge transfer and its benefits, we measured standard clustering metrics as an intrinsic evaluation of acoustic representation, and to emphasize applicability, we performed isolated-word recognition tasks using an English benchmark and a Cameroonian language dataset in low resources settings . Results demonstrate that acoustic representations consistently benefit from linguistic knowledge propagated through the graph. To our knowledge, this is the first demonstration of explicit cross-modal knowledge transfer for acoustic representation learning using HGNNs, highlighting a promising direction for speech representation in low-resource settings.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Yannick Yomie Nzeuhang, Paulin Melatagia Yonta, Marie Tahon. 2026-09-19. Enhancing speech representation learning with cross-modal knowledge transfer with HGNN under low resource settings: the case study of Yemba. https://arxiv.org/abs/2609.23194
Cite the original work for its findings. Save a collection to share your selection of sources.