arXiv Science⌕ Search

arXiv · 2610.08844

UZH-CL at ArA-DF 2026: Prompt-Tuned Foundation Models and Track-Adaptive Score Fusion for Arabic Speech Deepfake Detection

Abstract

Detecting synthetic and voice-converted speech remains difficult for low-resource languages with dialectal diversity, where systems must generalize across regional dialects and unseen acoustic channels. We present the UZH-CL submission to the ArA-DF 2026 Shared Task on Arabic speech deepfake detection, covering Track~1 (dialect generalization) and Track~2 (acoustic robustness). We freeze a W2V-BERT-2.0 backbone and adapt it with \emph{Wavelet Prompt Tuning}, updating under 1% of parameters, and aggregate multi-layer representations with cross-layer attention and a general attentive-statistics pooling head rather than a specialized graph backend. Complementary detectors are obtained by varying adaptation strategy, training data, augmentation, and encoder family. We find that the two shift types require different fusion regimes: a broad multi-window ensemble for dialect generalization, and a compact, channel-matched, center-crop ensemble for acoustic robustness. Official evaluation yields 1.96% EER on Track~1 (6th place) and 1.04% EER on Track~2 (3rd place), corresponding to 87% and 96% relative reductions over the XLS-R+AASIST baselines.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Aref Farhadipour, Teodora Vukovic, Petr Motlicek. 2026-10-01. UZH-CL at ArA-DF 2026: Prompt-Tuned Foundation Models and Track-Adaptive Score Fusion for Arabic Speech Deepfake Detection. https://arxiv.org/abs/2610.08844

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Towards Automated Clinical Behavioral Coding with Large Language Models: A Case study Using BOSCC recordings of Children

Autism spectrum disorder (ASD) is a neurodevelopmental condition characterized by differences in social communication and by restricted interests and repetitive behaviors. Treatment interventions often target social-communication skills, creating a need for reliable measures of behavioral change. The Brief Observation of Social Communication Change (BOSCC) is a validated treatment-response measure based on brief play and social-communication interactions between a child and trained examiner. The BOSCC coding process is resource-intensive and requires trained experts, motivating the automation of coding in order to improve scalability and accessibility. In this work, we evaluate general-purpose large language models (LLMs) for predicting speech-related BOSCC codes from different input representations. We compare transcript, diarized-transcript, and targeted audio conditions across 163 in-house recordings. LLMs are able to perform well in assessing verbal exchange, but do not perform as well when identifying atypical speech patterns. Additionally, performance varies considerably across scoring decisions, with no consistent pattern across diagnosis groups. An audit of model predictions indicates that applying the BOSCC coding criteria and interpreting ambiguous speech evidence remain challenges.

eess.AS↗

SmoothConv and DuplexConv: Complementary Mandarin Multi-Party Conversational Speech Corpora for Speech Interaction

Recent advances in large audio language models (LALMs) have driven the development of natural and intelligent speech interaction systems. Such systems need to model complex conversational behaviors, including turn-taking, overlapping speech, and speaker coordination. Multi-party conversations provide a realistic setting for studying these behaviors, yet existing Mandarin conversational corpora often lack synchronized participant-level speech tracks and comprehensive annotations. In this work, we introduce SmoothConv and DuplexConv, two complementary Mandarin multi-party conversational speech corpora totaling 2,100 hours. SmoothConv provides human-verified conversations for reliable analysis and evaluation, while DuplexConv offers large-scale automatically annotated conversations through a scalable pipeline for model training. Both corpora provide synchronized participant-level speech tracks and multi-dimensional fine-grained annotations. We further release the SmoothConv Benchmark and evaluate these resources on speech separation, multi-speaker automatic speech recognition (MSASR), and turn detection tasks. Experimental results demonstrate the utility of the proposed resources for multi-party speech interaction modeling. The datasets, benchmark, and related resources are publicly available.

eess.AS↗

Randomized Scores and Diverse Timbres: Augmenting Automatic Music Transcription with Online-Generated Data

Automatic music transcription (AMT) is limited by the scarcity of audio recordings paired with precise symbolic annotations. Synthetic data can provide supervision at scale, but it remains unclear whether effective transfer depends on realistic score structure or broad timbral coverage. We study these factors separately through an online sampler--renderer pipeline. A unified corruption sampler ranges from unmodified MIDI clips through partial corruption to deeply randomized note-event distributions. The renderer converts these events to audio while independently controlling instrument and timbral coverage. A fixed transcription model is trained jointly on offline recordings and newly rendered examples. Controlled ablations reveal an asymmetry between the two factors: moderate corruption of the note-event distribution does not impair transfer and can improve it, whereas broader renderer-side timbral support consistently improves out-of-domain generalization under a fixed note-event distribution. Finally, online-rendered examples complement real and existing synthetic data in a strong combined-data regime. These results suggest that synthetic AMT data should prioritize coverage of note-level attributes and their timbral realizations over realistic joint score structure.

eess.AS↗