arXiv · 2609.25948
Challenges of Multi-Speaker Extraction for Real Conversational Speech Enhancement
Abstract
Target-speaker and multi-speaker extraction are techniques for extracting speech from a desired speaker or desired speakers in the presence of other speakers and/or noise. Neural network approaches for this task are often trained and evaluated using simulated datasets, with balanced amounts of target speech and speaker enrolment samples which closely match the target speech. However, in real multi-party conversations, participants are often silent for more time than they are speaking, and their enrolment speech samples can differ substantially from the target speech in the conversation. These factors can impact the training and evaluation of these techniques on recordings of real conversations. This work proposes a new loss function, which helps mitigate the effect of excess silence in training, improving STOI from 0.55 to 0.60, and frequency-weighted segmental SNR from 4.35 to 5.12. Additionally, the impact of the mismatch between the enrolment speech and target speech is explored.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Robert Sutherland, Stefan Goetze, Jon Barker. 2026-09-22. Challenges of Multi-Speaker Extraction for Real Conversational Speech Enhancement. https://arxiv.org/abs/2609.25948
Cite the original work for its findings. Save a collection to share your selection of sources.