Indic-CLAP: Text--Audio Understanding in Indic Languages via Cross-Lingual Distillation
Contrastive Language--Audio Pretraining (CLAP) learns a joint text--audio representation space, which enables zero-shot audio understanding from natural-language descriptions. However, its English-only text encoder restricts CLAP to tasks and datasets expressed in English. We present Indic-CLAP, a multilingual text--audio model that extends CLAP to nine Indic languages spoken by over a billion people. Exploiting CLAP's dual-encoder structure, we propose to train only an Indic text encoder by distillation while keeping the audio encoder frozen, using machine-translated AudioCaps captions. We compare distillation from CLAP's text encoder, audio encoder, and a novel hybrid objective combining both. Our evaluation shows that the Indic-CLAP encoder inherits the cross-modal alignment from the CLAP teacher, as evidenced by the performance on cross-modal retrieval and zero-shot classification tasks. Further, we study the cross-modal alignment using modality gap analysis, which explains the performance trends.