arXiv · 2607.25085
Text-Prompted CLAP: Learning Text-Conditioned Audio Representations via Contrastive Learning
Abstract
Contrastive Language-Audio Pretraining (CLAP) learns aligned text and audio representations in a shared embedding space. However, independent encoding of each modality limits its ability to model cross-modal semantics in complex audio understanding and retrieval tasks. To address this limitation, this paper proposes Text-Prompted CLAP (TP-CLAP), a parameter-efficient extension of CLAP that introduces a cross-attention-based fusion module to incorporate textual prompts into audio features. TP-CLAP is trained using an audio multiple-choice question answering (AMCQA) framework, where it learns to align text-conditioned audio representations with text embeddings of correct answer choices via contrastive learning. Experiments demonstrate that TP-CLAP performs competitively with substantially larger audio-LLMs on audio question answering, while also improving the base CLAP model on conventional audio-text retrieval and zero-shot classification benchmarks. The learned representations are further fine-tuned for attribute-focused audio-to-audio retrieval, showing that TP-CLAP consistently outperforms the standard CLAP baseline in music retrieval tasks.
Explore related subjects
Keep this discovery
Mohan Li, Rama Doddipatla, Philip C. Woodland. 2026-07-27. Text-Prompted CLAP: Learning Text-Conditioned Audio Representations via Contrastive Learning. https://arxiv.org/abs/2607.25085
Cite the original work for its findings. Save a collection to share your selection of sources.