arXiv · 2610.04180
MOV-AAD: A Large-Scale Multimodal Dataset for Auditory Attention Decoding During Moving Conversations
Abstract
Auditory attention decoding (AAD) is often evaluated on static, simplified speech scenes that poorly match everyday listening. We introduce MOV-AAD, a large-scale dataset for studying auditory attention under moving, naturalistic conversations. MOV-AAD combines 64-channel EEG with synchronized physiological recordings, including eye tracking, respiration, galvanic skin response, heart rate, peripheral oxygen saturation, body temperature, body motion, and photoplethysmography, enabling analysis of cross-modal neural and physiological markers of attention and listening effort. This dataset uses a more ecologically valid listening paradigm with dynamically moving conversational speech sources and behavioral measures of attentional engagement. MOV-AAD supports research on robust AAD in realistic spatial dynamics, multimodal attention modeling, listening effort, and intersubject neural responses, providing a resource for benchmarking selective auditory attention in naturalistic listening.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Xiaomin He, Vishal Choudhari, Tristan J. Spratt, Aarya Raghavan, Richard T. Lee, Nima Mesgarani. 2026-10-03. MOV-AAD: A Large-Scale Multimodal Dataset for Auditory Attention Decoding During Moving Conversations. https://doi.org/10.21437/interspeech.2026-3556
Cite the original work for its findings. Save a collection to share your selection of sources.