arXiv · 2609.23479
Generative Learning for Ambisonic Upscaling
Abstract
Ambisonics Upscaling (AU) aims to enhance the spatial resolution of sound fields by estimating high-order Ambisonics (HOA) components from low-order observations. While deep learning and model-based strategies have been considered for AU, both approaches exhibit significant performance degradation in realistic scenarios, where reverberant sound fields violate the directional sparsity inherent to discriminative mappings. In this work, we address AU as a generative task rather than a deterministic reconstruction, expanding generative modeling to specifically target the recovery of spatial information in reverberant speech. We investigate two dominant continuous-time generative paradigms, adapting both Score-based Generative Model and Flow Matching to these complex acoustic settings. We provide an extensive numerical study comparing our methods against state-of-the-art baselines in various acoustic scenarios. Additionally, we conduct subjective listening tests to evaluate the perceived quality and spatial accuracy of the proposed generative framework across various reverberant scenarios. The studies reveal that Flow Matching consistently outperforms both its discriminative counterparts and Diffusion-based paradigms in all reverberant settings.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Amit Milstein, Nir Shlezinger, Boaz Rafaely. 2026-09-20. Generative Learning for Ambisonic Upscaling. https://arxiv.org/abs/2609.23479
Cite the original work for its findings. Save a collection to share your selection of sources.