arXiv · 2610.08847
Timestamped Hindi speech transcription using Whisper
Abstract
Time accurate speech transcription is essential in critical appli- cations, such as medical conversation transcription, reading as- sessment, and atypical speech recognition. Encoder-decoder auto- matic speech recognition models, such as Whisper, do not produce word-level timestamps out-of-the-box for all languages. Existing approaches to word-level time stamp estimation either use an exter- nal CTC model capable of frame level phoneme/character prediction which is computationally expensive, or use the cross attention scores internal to the model which gives coarse time stamps. We consider timestamped speech transcription for Hindi in this paper. We pro- pose a CTC-based approach, in which we train a low-complexity character level CTC head directly on the Whisper encoder output, to address the complexity and granularity limitations of existing works. We evaluate the system using manually-annotated word-level times- tamps. Experiments show that the proposed internal-CTC based approach performs similar to using an external CTC model and better than the cross attention score-based approaches. Further, we study the usage of CTC-head output to detect and mitigate hallu- cinations in conversational speech. Our results show that a simple token length criterion derived from the CTC output can improve the transcription quality for short segments where hallucinations are more probable.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Sanat Kumar Agrawal, Srikanth Raj Chetupalli. 2026-10-01. Timestamped Hindi speech transcription using Whisper. https://arxiv.org/abs/2610.08847
Cite the original work for its findings. Save a collection to share your selection of sources.