arXiv · 2609.15218
Word Timestamps and Speaker Attribution with a Non-Autoregressive LLM
Abstract
Timestamps and speaker attribution are useful additions to speech recognition, creating a rich text transcript. This information can either be extracted during transcription or aligned to a given transcript. In this paper we present models that add timestamps and speaker information to a given transcript using a non-autoregressive LLM-based architecture. Compared to an autoregressive model built from similar components, the models are more accurate and annotate a given transcript one to two orders of magnitude faster. Compared to other models, our models achieve state-of-the-art timestamp accuracy and the best cpWER for speaker attribution.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Zvi Kons, Avihu Dekel, Hagai Aronowitz, Vishal Sunder, Ron Hoory. 2026-09-14. Word Timestamps and Speaker Attribution with a Non-Autoregressive LLM. https://arxiv.org/abs/2609.15218
Cite the original work for its findings. Save a collection to share your selection of sources.