arXiv Science⌕ Search

arXiv · 2609.39162

Training-Free Affinity Fusion of Neural and Embedding-Based Speaker Diarization

Abstract

Speaker diarization systems based on speaker embeddings and neural diarization exploit complementary forms of speaker information, but their intermediate representations are not directly compatible. We introduce Training-Free Affinity Fusion (TFAF), which integrates the speaker structure inferred by a neural diarizer into an embedding-based diarization system. The neural speaker partition is used to condition local speaker representations, from which we construct a continuous affinity matrix and combine it with the embedding-based acoustic affinity before a single global clustering step. The method requires no additional training, shared embedding space, speaker-label alignment, or hard transfer of the neural diarizer's speaker count. Experiments on AMI and CALLHOME show consistent DER improvements over both constituent systems; on AMI, fusion also improves speaker-attributed transcription. Ablations show that the neural speaker partition accounts for most of the gain, while retaining the continuous embedding-based affinities provides additional benefit over hard partition fusion.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Yehoshua Dissen, Joseph Keshet, Eduard Golshtein. 2026-09-30. Training-Free Affinity Fusion of Neural and Embedding-Based Speaker Diarization. https://arxiv.org/abs/2609.39162

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

WASIL: In-the-Wild Arabic Spoken Interactions with LLMs

Large Language Models (LLMs) voice assistants are commonly built as cascaded Automatic Speech recognition (ASR) to LLM systems, where recognition errors can distort user intent. Dislikes may also arise from ambiguous, out-of-domain, or non-request turns, making it hard to isolate ASR effects. We release WASIL (it denotes connection or linking in Arabic): in-the-wild Arabic spoken interaction prompts with audio, ASR hypotheses, assistant responses, and explicit like/dislike feedback (8,529 turns; 14.2% dislikes), plus a 2,000-turn test set covering Modern Standard Arabic (MSA) and four major dialects with their labels. We provide low-cost gold transcripts via multi-ASR agreement-guided post-editing and annotate answerability (answerable, ambiguous/needs-clarification, unsupported, not-a-request/noise) to separate intrinsic unanswerability from ASR-induced degradation. Finally, we describe scalable reference-free evaluation of responses from ASR vs. gold transcripts using multi-judge LLM scoring.

cs.SD↗

Trigger Sound Suppression for Misophonia

Misophonia, a disorder of decreased tolerance to specific sounds, affects 5-20% of the population, yet sufferers have no good options: therapy helps a minority, and earplugs or noise cancellation silence everything. We present a study for neural trigger sound suppression for misophonia, selectively removing trigger sounds. We curate a dataset covering the 10 most common trigger classes. Using streaming dual-path networks operating on 6 ms audio chunks, we explore both one-hot and multi-hot-conditioned models that suppress 1-3 triggers from the acoustic scene. We validate our model outputs in a listening study with 30 adults with clinically elevated misophonia impairment. Participants reported significantly lower distress and arousal, and improved valence, for suppressed audio.

cs.SD↗

Bad: Taming the Bioacoustic Data Deluge with a Bat Activity Detector

Passive Acoustic Monitoring of bats generates massive ultrasonic datasets (>27 GB/night per node), straining edge storage and battery life. Legacy triggers fail against acoustic confusers, while deep models exceed microcontroller limits. We present a hardware-aware Bat Activity Detector (BAD) specifically designed to discriminate bat calls from hard biological and environmental confusers across variable sampling rates (192-384 kHz). Tailored for the Silicon Labs EFM32PG26 (MVP) in 8-bit integer precision, our model achieves 100 percent hardware offload across all 14 layers (17.2 KB Flash, 73.1 KB RAM). End-to-end preprocessing (74.00 ms for 76 frames) and inference (30.00 ms) of 100 ms clips at 192 kHz require 104.00 ms per clip. On spatially out-of-domain recordings under a realistic low-prevalence regime (r_pos = 0.05), BAD achieves an AUC-ROC of 0.9748 and suppresses 99.4% of non-target noise frames while retaining 65.3% of bat calls - delivering a >33x precision gain over classical Goertzel baselines.

cs.SD↗