arXiv · 2609.23152
Bearings: Self-Supervised Soundfield Embeddings from First-Order Ambisonics
Abstract
Recently proposed self-supervised audio encoders learn powerful general-purpose representations of sound scenes, yet they are spatially blind. To supply the missing spatial representation of sound scenes, we introduce Bearings. Bearings is a self-supervised framework that learns soundfield embeddings from unlabeled first-order Ambisonics. We pre-train a masked auto-encoder paired with a decoder conditioned on frozen acoustic embeddings from an off-the-shelf single-channel audio encoder. Our results show that the resulting soundfield embeddings form a reusable stream that can be attached to frozen acoustic encoders with a lightweight trainable fusion head. On sound event localization and detection, concatenating our soundfield embeddings with acoustic representations provides the missing spatial information and enables joint detection and localization, raising the location-dependent F-score from below 4 to 50 on TAU-NIGENS 2021 and 39 on STARSS23. To our knowledge, Bearings is the first self-supervised soundfield encoder whose embeddings plug into frozen acoustic encoders without retraining either model.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Goksenin Yuksel, Marcel van Gerven, Kiki van der Heijden. 2026-09-19. Bearings: Self-Supervised Soundfield Embeddings from First-Order Ambisonics. https://arxiv.org/abs/2609.23152
Cite the original work for its findings. Save a collection to share your selection of sources.