arXiv · 2609.26648
ROAM-ASD: Robust Open-World Active Speaker Detection with Flexible Multimodal Fusion
Abstract
Active speaker detection (ASD) requires reliable association between visible faces and acoustic speech, yet existing systems often degrade under challenging domains or incomplete observations. We introduce ROAM-ASD, a robust audiovisual framework that jointly models audio, full-face, and fine-grained mouth representations. A unified joint self-attention mechanism processes all input streams together with modality-agnostic query tokens, enabling direct interaction among available modality inputs. Modality dropout further improves robustness when input streams are unavailable. ROAM-ASD achieves state-of-the-art performance across five ASD benchmarks: 98.8% mAP on WASD, 87.9% on UniTalk, 96.5% on AVA, 99.3% on ASW, and 98.2% on Talkies, improving over previous best systems by 5.1, 4.7, 0.9, 1.0, and 2.1 mAP points, respectively. ROAM-ASD also substantially improves zero-shot cross-dataset generalization and remains robust to missing observations.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Pu Wang, Yujun Wang, Hugo Van hamme. 2026-09-22. ROAM-ASD: Robust Open-World Active Speaker Detection with Flexible Multimodal Fusion. https://arxiv.org/abs/2609.26648
Cite the original work for its findings. Save a collection to share your selection of sources.