arXiv Science⌕ Search

arXiv · 2610.11272

Using the enrollment gallery as evidence: affine-invariant score calibration for target speaker tagging

Abstract

Target speaker tagging (TST) assigns enrolled speaker identities to diarized segments of multi-speaker recordings. The enrollment utterances of the other registered speakers are an attractive source of evidence for calibrating each verification decision: they share the deployment domain of the test material and require no external cohort set. We show, however, that they cannot serve as the cohort of conventional score normalization. Enrollments tend to cluster by recording session, so the per-speaker cohort statistics reflect enrollment proximity to the rest of the gallery rather than impostor behavior, and normalization then rejects entire speakers. We propose gallery affinity verification, a cohort-free calibration that exploits the gallery through two affine-invariant terms, score dispersion and enrollment-profile deviance, and is therefore immune to such per-speaker shifts. On a synthetic benchmark and an in-house meeting corpus, it improves tagging accuracy with and without conventional score normalization and adds further gains when combined with it.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Hee-Soo Heo, Minjae Lee, Youngki Kwon, Bong-Jin Lee. 2026-10-08. Using the enrollment gallery as evidence: affine-invariant score calibration for target speaker tagging. https://arxiv.org/abs/2610.11272

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Towards Automated Clinical Behavioral Coding with Large Language Models: A Case study Using BOSCC recordings of Children

Autism spectrum disorder (ASD) is a neurodevelopmental condition characterized by differences in social communication and by restricted interests and repetitive behaviors. Treatment interventions often target social-communication skills, creating a need for reliable measures of behavioral change. The Brief Observation of Social Communication Change (BOSCC) is a validated treatment-response measure based on brief play and social-communication interactions between a child and trained examiner. The BOSCC coding process is resource-intensive and requires trained experts, motivating the automation of coding in order to improve scalability and accessibility. In this work, we evaluate general-purpose large language models (LLMs) for predicting speech-related BOSCC codes from different input representations. We compare transcript, diarized-transcript, and targeted audio conditions across 163 in-house recordings. LLMs are able to perform well in assessing verbal exchange, but do not perform as well when identifying atypical speech patterns. Additionally, performance varies considerably across scoring decisions, with no consistent pattern across diagnosis groups. An audit of model predictions indicates that applying the BOSCC coding criteria and interpreting ambiguous speech evidence remain challenges.

eess.AS↗

SmoothConv and DuplexConv: Complementary Mandarin Multi-Party Conversational Speech Corpora for Speech Interaction

Recent advances in large audio language models (LALMs) have driven the development of natural and intelligent speech interaction systems. Such systems need to model complex conversational behaviors, including turn-taking, overlapping speech, and speaker coordination. Multi-party conversations provide a realistic setting for studying these behaviors, yet existing Mandarin conversational corpora often lack synchronized participant-level speech tracks and comprehensive annotations. In this work, we introduce SmoothConv and DuplexConv, two complementary Mandarin multi-party conversational speech corpora totaling 2,100 hours. SmoothConv provides human-verified conversations for reliable analysis and evaluation, while DuplexConv offers large-scale automatically annotated conversations through a scalable pipeline for model training. Both corpora provide synchronized participant-level speech tracks and multi-dimensional fine-grained annotations. We further release the SmoothConv Benchmark and evaluate these resources on speech separation, multi-speaker automatic speech recognition (MSASR), and turn detection tasks. Experimental results demonstrate the utility of the proposed resources for multi-party speech interaction modeling. The datasets, benchmark, and related resources are publicly available.

eess.AS↗

Randomized Scores and Diverse Timbres: Augmenting Automatic Music Transcription with Online-Generated Data

Automatic music transcription (AMT) is limited by the scarcity of audio recordings paired with precise symbolic annotations. Synthetic data can provide supervision at scale, but it remains unclear whether effective transfer depends on realistic score structure or broad timbral coverage. We study these factors separately through an online sampler--renderer pipeline. A unified corruption sampler ranges from unmodified MIDI clips through partial corruption to deeply randomized note-event distributions. The renderer converts these events to audio while independently controlling instrument and timbral coverage. A fixed transcription model is trained jointly on offline recordings and newly rendered examples. Controlled ablations reveal an asymmetry between the two factors: moderate corruption of the note-event distribution does not impair transfer and can improve it, whereas broader renderer-side timbral support consistently improves out-of-domain generalization under a fixed note-event distribution. Finally, online-rendered examples complement real and existing synthetic data in a strong combined-data regime. These results suggest that synthetic AMT data should prioritize coverage of note-level attributes and their timbral realizations over realistic joint score structure.

eess.AS↗