arXiv Science⌕ Search

arXiv · 2610.08845

Unified Shared Encoder in Spoof-Aware Speaker Verification with Hybrid Wavelet Prompt Tuning

Abstract

Spoofing-aware speaker verification (SASV) must confirm who is speaking and that the speech is genuine, but current systems gain one at the expense of the other. Modular cascades are accurate yet run two large self-supervised encoders, whereas end-to-end models are efficient but lose accuracy when a single embedding has to serve two conflicting objectives. We show that this trade-off between efficiency and specialization is avoidable. A single frozen W2V-BERT~2.0 backbone is adapted per task by deep hybrid Wavelet Prompt Tuning (WPT), so each branch obtains its own view of the shared encoder through dedicated prompts and a task head, and the scores are combined only at inference. Training only about 7M parameters, a small fraction of those a strong two-encoder cascade requires, the system surpasses it on SpoofCeleb evaluation with 0.03% CM-EER and 0.038 min a-DCF against 0.16% and 0.047. Because the backbone is shared and frozen, new branches attach without retraining the heads and prompts already in place. On ASVspoof5, with adversarial attacks and codec distortions, it reaches 0.092 min a-DCF and 3.60% CM-EER using only the provided data, indicating that the design holds under realistic in-the-wild threats.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Aref Farhadipour, Srikanth Madikeri, Teodora Vukovic, Volker Dellwo, Petr Motlicek. 2026-10-01. Unified Shared Encoder in Spoof-Aware Speaker Verification with Hybrid Wavelet Prompt Tuning. https://arxiv.org/abs/2610.08845

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

WaveScat: Wavelet Scattering Front-Ends with Self-Supervised Features for Speech Deepfake Detection

Existing front-ends for speech deepfake detection are primarily categorized into two types. Hand-crafted filterbank features are transparent but limited in capturing higher-level information. SSL features, in turn, lack interpretability and may overlook fine-grained spectral anomalies. We propose WaveScat, a novel family of feature extractors that combines the best of both worlds via the wavelet scattering transform (WST), which cascades wavelet convolutions with modulus nonlinearities to produce deformation-stable, multi-scale features. Experiments on the recent Deepfake-Eval-2024 benchmark, together with cross-dataset evaluations on SpoofCeleb, In-the-Wild, and ASVspoof 5, show that WaveScat outperforms existing front-ends by a wide margin. Our analysis reveals that a small averaging scale combined with high-frequency and directional resolutions is critical for capturing subtle artifacts. This underscores the value of stable and translation-invariant features for speech deepfake detection. The code and supplementary materials are available at https://github.com/xxuan-acoustics/WaveScat.

eess.AS↗

Do Multimodal Large Language Models Need Reasoning to Classify Dementia from Speech?

Multimodal large language models (MLLMs) have emerged as a promising approach for improving the accuracy, transferability, and explainability of automatic dementia classification (ADC) systems from voice recordings. Yet it remains unclear whether their reasoning capabilities are beneficial for ADC, and how such capabilities should be leveraged. In this paper, we conduct a careful evaluation of reasoning MLLMs for ADC and show that naive strategies, such as relying on text-based rationales, can lead to hallucinated and inconsistent rationales for diagnosis and yield inferior ADC performance compared with LLM-free baselines. To overcome this limitation, we propose \textbf{De}mentia \textbf{T}hinker with Nonlinear \textbf{A}daptor and Re\textbf{i}nforcement \textbf{L}earning (DeTAiL), an adaptor-based framework that exploits the internal representations of reasoning MLLMs for improved dementia classification. Across two dementia datasets with distinct test formats and label granularities, DeTAiL consistently outperforms strong baselines and methods that rely on text-based rationales. Code and demo will be released upon acceptance.

eess.AS↗

Beyond Words: Towards Effective Modeling of Non-Verbal Vocalizations in ASR

Modern automatic speech recognition (ASR) systems excel at transcribing lexical content but often omit nonverbal vocalizations (NVs), such as laughter, breaths, coughs, and cries, that carry conversational and affective information. Modeling NVs in ASR is challenging because NV annotations are sparse and highly long-tailed, with frequent categories such as breaths and laughter dominating rarer events such as cries and coughs. We study three data-centric strategies for improving low-resource NV recognition: (1) a two-stage curriculum that first maps all NV events to a generic token and then fine-tunes on target categories; (2) inter-token transfer from high-resource events, such as laughter and breath, to rare events, such as crying; and (3) voice-conversion augmentation with class balancing. Experiments show that shared acoustic structure across vocal events can be exploited to improve rare-category detection while preserving lexical ASR quality.

eess.AS↗