arXiv · 2609.23114
Speech Language Models for Full-Meeting Speaker Diarization: Capabilities and Limitations
Abstract
Recent advances in Speech Language Models (SpeechLMs), which integrate large language models with speech foundation models, have enabled unified sequence modeling of speech processing tasks. However, many SpeechLM-based approaches to speaker diarization (SD) are tightly coupled with automatic speech recognition (ASR) and evaluated using word-level metrics, making it difficult to assess SD performance independent of ASR accuracy. In this work, we investigate ESPnet-SpeechLM as a token-based backbone for generating SD hypotheses, formulating SD as autoregressive generation of structured tokens conditioned on acoustic input. We systematically compare two output representations: an event-based representation that explicitly models speaker turn onset and offset timestamps, and a frame-based representation that predicts frame-level speaker activity. To provide structured conversational cues, we further incorporate auxiliary tasks including speech activity detection, overlapped speech detection, and speaker turn counting within the output sequence. Across multiple meeting datasets, we find that event-based representations produce more stable and consistent SD outputs than frame-based representations. Our analysis shows that outputs generated by SpeechLMs encode useful temporal SD structure, but full-meeting SD remains limited by recording-level speaker tracking and overlap-related misses. Explicit speaker-linking post-processing substantially reduces speaker confusion, suggesting that robust SpeechLM-based SD requires persistent speaker tracking and overlap-aware generation.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Jialu Li, Jinchuan Tian, Shinji Watanabe. 2026-09-19. Speech Language Models for Full-Meeting Speaker Diarization: Capabilities and Limitations. https://arxiv.org/abs/2609.23114
Cite the original work for its findings. Save a collection to share your selection of sources.