arXiv · 2609.31898
MAESTRO: a Multimodal Auditory-attention Egocentric Speech-TRacking Open corpus
Abstract
Humans rely on gaze, head movements, and visual cues to attend to speakers in noisy environments, yet auditory attention decoding (AAD) has been studied primarily using electroencephalography (EEG). We introduce the Multimodal Auditory-attention Egocentric Speech-TRacking Open (MAESTRO) corpus, the first AAD dataset to simultaneously record EEG, eye gaze, pupillometry, egocentric video, and head inertial measurement unit (IMU) data. MAESTRO includes four competing speakers and background noise across multiple signal-to-noise ratio (SNR) conditions, enabling attention decoding under realistic listening scenarios. Through a four-speaker attention decoding benchmark, we show that combining behavioral and physiological signals improves decoding performance over EEG-only approaches, enabling future advances in multimodal auditory attention decoding. These findings open the door to new applications, analyses, and methodological advances in multimodal AAD. The complete dataset is publicly available at https://huggingface.co/datasets/aspire-osu/maestro-eeg-dataset . The official code repository is available at https://github.com/ASPIRE-OSU/MAESTRO .
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
K M Naimul Hassan, Ali Alavi, Donald S. Williamson. 2026-09-25. MAESTRO: a Multimodal Auditory-attention Egocentric Speech-TRacking Open corpus. https://arxiv.org/abs/2609.31898
Cite the original work for its findings. Save a collection to share your selection of sources.