arXiv · 2609.39162
Training-Free Affinity Fusion of Neural and Embedding-Based Speaker Diarization
Abstract
Speaker diarization systems based on speaker embeddings and neural diarization exploit complementary forms of speaker information, but their intermediate representations are not directly compatible. We introduce Training-Free Affinity Fusion (TFAF), which integrates the speaker structure inferred by a neural diarizer into an embedding-based diarization system. The neural speaker partition is used to condition local speaker representations, from which we construct a continuous affinity matrix and combine it with the embedding-based acoustic affinity before a single global clustering step. The method requires no additional training, shared embedding space, speaker-label alignment, or hard transfer of the neural diarizer's speaker count. Experiments on AMI and CALLHOME show consistent DER improvements over both constituent systems; on AMI, fusion also improves speaker-attributed transcription. Ablations show that the neural speaker partition accounts for most of the gain, while retaining the continuous embedding-based affinities provides additional benefit over hard partition fusion.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Yehoshua Dissen, Joseph Keshet, Eduard Golshtein. 2026-09-30. Training-Free Affinity Fusion of Neural and Embedding-Based Speaker Diarization. https://arxiv.org/abs/2609.39162
Cite the original work for its findings. Save a collection to share your selection of sources.