arXiv · 2608.28661
Neural Multichannel Distant Speaker Diarization and Source Separation with Beta Speaker Activity Prior
Abstract
Distant speaker diarization remains challenging due to adverse acoustic conditions, varying numbers of speakers and overlapping speech. While data-driven approaches have shown strong performance, model-driven methods offer a compelling alternative by leveraging spatial information from multichannel recordings. This paper is motivated to propose a Bayesian diarization model for a model-driven method called neural FCASA to enhance its robustness. Specifically, we propose a beta prior over speaker activity and hence a variational lower bound objective that can be seen as a regularized continuous speaker activity score in place of the original cross-entropy loss to train the diarization model. Our experiments show significant improvements in terms of Diarization Error Rate by at least 3% (16% relatively) and Jaccard Error Rate by at least 4% (20% relatively) on the AMI dataset compared to the baseline.
Explore related subjects
Keep this discovery
Sicheng Mao, Mathieu Fontaine, Anthony Larcher, Roland Badeau. 2026-08-21. Neural Multichannel Distant Speaker Diarization and Source Separation with Beta Speaker Activity Prior. https://arxiv.org/abs/2608.28661
Cite the original work for its findings. Save a collection to share your selection of sources.
Discover connections
Connections use source metadata and explicit phrase matches, not verified experimental comparisons.