arXiv · 2610.08844
UZH-CL at ArA-DF 2026: Prompt-Tuned Foundation Models and Track-Adaptive Score Fusion for Arabic Speech Deepfake Detection
Abstract
Detecting synthetic and voice-converted speech remains difficult for low-resource languages with dialectal diversity, where systems must generalize across regional dialects and unseen acoustic channels. We present the UZH-CL submission to the ArA-DF 2026 Shared Task on Arabic speech deepfake detection, covering Track~1 (dialect generalization) and Track~2 (acoustic robustness). We freeze a W2V-BERT-2.0 backbone and adapt it with \emph{Wavelet Prompt Tuning}, updating under 1% of parameters, and aggregate multi-layer representations with cross-layer attention and a general attentive-statistics pooling head rather than a specialized graph backend. Complementary detectors are obtained by varying adaptation strategy, training data, augmentation, and encoder family. We find that the two shift types require different fusion regimes: a broad multi-window ensemble for dialect generalization, and a compact, channel-matched, center-crop ensemble for acoustic robustness. Official evaluation yields 1.96% EER on Track~1 (6th place) and 1.04% EER on Track~2 (3rd place), corresponding to 87% and 96% relative reductions over the XLS-R+AASIST baselines.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Aref Farhadipour, Teodora Vukovic, Petr Motlicek. 2026-10-01. UZH-CL at ArA-DF 2026: Prompt-Tuned Foundation Models and Track-Adaptive Score Fusion for Arabic Speech Deepfake Detection. https://arxiv.org/abs/2610.08844
Cite the original work for its findings. Save a collection to share your selection of sources.