arXiv · 2610.08845
Unified Shared Encoder in Spoof-Aware Speaker Verification with Hybrid Wavelet Prompt Tuning
Abstract
Spoofing-aware speaker verification (SASV) must confirm who is speaking and that the speech is genuine, but current systems gain one at the expense of the other. Modular cascades are accurate yet run two large self-supervised encoders, whereas end-to-end models are efficient but lose accuracy when a single embedding has to serve two conflicting objectives. We show that this trade-off between efficiency and specialization is avoidable. A single frozen W2V-BERT~2.0 backbone is adapted per task by deep hybrid Wavelet Prompt Tuning (WPT), so each branch obtains its own view of the shared encoder through dedicated prompts and a task head, and the scores are combined only at inference. Training only about 7M parameters, a small fraction of those a strong two-encoder cascade requires, the system surpasses it on SpoofCeleb evaluation with 0.03% CM-EER and 0.038 min a-DCF against 0.16% and 0.047. Because the backbone is shared and frozen, new branches attach without retraining the heads and prompts already in place. On ASVspoof5, with adversarial attacks and codec distortions, it reaches 0.092 min a-DCF and 3.60% CM-EER using only the provided data, indicating that the design holds under realistic in-the-wild threats.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Aref Farhadipour, Srikanth Madikeri, Teodora Vukovic, Volker Dellwo, Petr Motlicek. 2026-10-01. Unified Shared Encoder in Spoof-Aware Speaker Verification with Hybrid Wavelet Prompt Tuning. https://arxiv.org/abs/2610.08845
Cite the original work for its findings. Save a collection to share your selection of sources.