arXiv Science⌕ Search

arXiv subjects

Sanat Kumar Agrawal

Publications and source records attributed to Sanat Kumar Agrawal.

2 recordsLinked to original sources

Indic-CLAP: Text--Audio Understanding in Indic Languages via Cross-Lingual Distillation

Contrastive Language--Audio Pretraining (CLAP) learns a joint text--audio representation space, which enables zero-shot audio understanding from natural-language descriptions. However, its English-only text encoder restricts CLAP to tasks and datasets expressed in English. We present Indic-CLAP, a multilingual text--audio model that extends CLAP to nine Indic languages spoken by over a billion people. Exploiting CLAP's dual-encoder structure, we propose to train only an Indic text encoder by distillation while keeping the audio encoder frozen, using machine-translated AudioCaps captions. We compare distillation from CLAP's text encoder, audio encoder, and a novel hybrid objective combining both. Our evaluation shows that the Indic-CLAP encoder inherits the cross-modal alignment from the CLAP teacher, as evidenced by the performance on cross-modal retrieval and zero-shot classification tasks. Further, we study the cross-modal alignment using modality gap analysis, which explains the performance trends.

eess.AS↗

Timestamped Hindi speech transcription using Whisper

Time accurate speech transcription is essential in critical appli- cations, such as medical conversation transcription, reading as- sessment, and atypical speech recognition. Encoder-decoder auto- matic speech recognition models, such as Whisper, do not produce word-level timestamps out-of-the-box for all languages. Existing approaches to word-level time stamp estimation either use an exter- nal CTC model capable of frame level phoneme/character prediction which is computationally expensive, or use the cross attention scores internal to the model which gives coarse time stamps. We consider timestamped speech transcription for Hindi in this paper. We pro- pose a CTC-based approach, in which we train a low-complexity character level CTC head directly on the Whisper encoder output, to address the complexity and granularity limitations of existing works. We evaluate the system using manually-annotated word-level times- tamps. Experiments show that the proposed internal-CTC based approach performs similar to using an external CTC model and better than the cross attention score-based approaches. Further, we study the usage of CTC-head output to detect and mitigate hallu- cinations in conversational speech. Our results show that a simple token length criterion derived from the CTC output can improve the transcription quality for short segments where hallucinations are more probable.

eess.AS↗