arXiv Science⌕ Search

arXiv · 2610.05055

Influence of Geometrical Acoustic Simulator Complexity on a Trained Multisource Localizer

Abstract

Learning-based acoustic sound source localization and detection (SSLD) requires large labeled datasets covering diverse acoustic conditions. Since obtaining measured data is costly, training commonly uses simulated data, while practical devices must operate under real-world conditions. However, higher simulation complexity increases data-generation cost, and the complexity required for reliable generalization remains unclear. In practice, SSLD methods often rely on efficient geometrical acoustic simulation, typically the image-source method. This work investigates how geometrical acoustic simulation complexity affects the real-world performance of a common multi-source SSLD model. We train the model with simulators ranging from anechoic conditions to high-order image-source simulations, optionally including diffuse reverberation, array simulation, and randomized image-source positions, and evaluate them on three measurement-based test datasets. Results show that anechoic simulation is insufficient, while medium-complexity image-source simulations already provide strong real-world performance. Further increases in complexity yield only marginal gains, with the best performance obtained using the highest tested image-source order, array simulation, and randomized image-source positions. For this configuration, measured-domain performance approaches within-domain simulated performance, suggesting limited benefit from further increasing simulation complexity. These findings show how the complexity-performance trade-off can be exploited in future data-driven SSLD, and highlight image-source randomization as an efficient way to improve generalization.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Fabian Staub, Nils Meyer-Kahlen, Thomas Deppisch, Sergio de las Heras, Florian Klein, Stephan Werner, Johannes M. Arend. 2026-10-04. Influence of Geometrical Acoustic Simulator Complexity on a Trained Multisource Localizer. https://arxiv.org/abs/2610.05055

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Personal VAD: Speaker-Conditioned Voice Activity Detection

In this paper, we propose "personal VAD", a system to detect the voice activity of a target speaker at the frame level. This system is useful for gating the inputs to a streaming on-device speech recognition system, such that it only triggers for the target user, which helps reduce the computational cost and battery consumption, especially in scenarios where a keyword detector is unpreferable. We achieve this by training a VAD-alike neural network that is conditioned on the target speaker embedding or the speaker verification score. For each frame, personal VAD outputs the probabilities for three classes: non-speech, target speaker speech, and non-target speaker speech. Under our optimal setup, we are able to train a model with only 130K parameters that outperforms a baseline system where individually trained standard VAD and speaker recognition networks are combined to perform the same task.

eess.AS↗

Textual Echo Cancellation

In this paper, we propose Textual Echo Cancellation (TEC) - a framework for cancelling the text-to-speech (TTS) playback echo from overlapping speech recordings. Such a system can largely improve speech recognition performance and user experience for intelligent devices such as smart speakers, as the user can talk to the device while the device is still playing the TTS signal responding to the previous query. We implement this system by using a novel sequence-to-sequence model with multi-source attention that takes both the microphone mixture signal and source text of the TTS playback as inputs, and predicts the enhanced audio. Experiments show that the textual information of the TTS playback is critical to enhancement performance. Besides, the text sequence is much smaller in size compared with the raw acoustic signal of the TTS playback, and can be immediately transmitted to the device or ASR server even before the playback is synthesized. Therefore, our proposed approach effectively reduces Internet communication and latency compared with alternative approaches such as acoustic echo cancellation (AEC).

eess.AS↗

Multiplexing Neural Audio Watermarks with Adaptive Routing

Audio watermarking supports speech authenticity verification. We study whether combining heterogeneous watermarks can retain at least one detectable provenance signal when their failure modes differ. This Any-survival objective concerns complementary evidence retention, not simultaneous survival of all constituent marks or recovery of all payloads. To our knowledge, this is the first systematic study of neural audio watermark multiplexing, with a scoped benchmark covering five released systems, parallel and sequential baselines, and 14 evaluation conditions. The benchmark shows complementary failure modes, but naive composition does not reliably turn them into system-level robustness under the Any-survival objective. We therefore formulate multiplexing as watermark allocation and study perceptual-adaptive time-frequency multiplexing (PA-TFM), a training-free routing method, and MaskNet, a learned time-domain router for separately trained systems with native detectors. We use the five-system scoped benchmark to characterize multiplexing behavior and the AudioSeal-PerTh pair as a representative heterogeneous pair for adaptive watermark routing. Compared with direct parallel composition, MaskNet improves average TPR@1%FPR from 0.75 to 0.88 and SNR from 15.20 dB to 25.36 dB; compared with PA-TFM, it gives a clear system-level robustness gain under the Any-survival objective while retaining similar high-fidelity behavior. A SpeechTokenizer-aware case study further improves average TPR@1%FPR to 0.91 and SpeechTokenizer robustness from 0.20 to 0.60, showing that channel-adapted constituents can enter the same routing framework.

eess.AS↗