arXiv · 2610.06691
Revisiting Label-Free Speaker Embedding Enhancement with vMF Profile Likelihood
Abstract
Embedding enhancement improves speaker verification under acoustic mismatch without modifying a frozen backbone. Recent work has established a practical label-free setting for this task, but often adopts increasingly structured formulations. Here, the clean target is directly observed during training, making enhancement a matching problem on the unit hypersphere. We model the clean target with a von Mises--Fisher (vMF) likelihood and profile out a sample-wise concentration parameter, yielding a simple closed-form objective with adaptive weighting. Across VoxCeleb1, VoxSRC23, CN-Celeb, VOiCES, and VC-Mix, the proposed method largely preserves the baseline and gives clearer gains on challenging mismatch sets. It also remains stable under a broad single-view recipe, where a recent diffusion baseline becomes less reliable in controlled comparisons. These results suggest that effective label-free embedding enhancement in this setting does not require a highly structured formulation.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Seunghwan Kim, Jinyong Kim, Sooyoung Yang, Youngjin Ko, Myungjoo Kang. 2026-10-05. Revisiting Label-Free Speaker Embedding Enhancement with vMF Profile Likelihood. https://doi.org/10.21437/interspeech.2026-1146
Cite the original work for its findings. Save a collection to share your selection of sources.