arXiv · 2303.08702
Beamformer-Guided Target Speaker Extraction
Abstract
We propose a Beamformer-guided Target Speaker Extraction (BG-TSE) method to extract a target speaker's voice from a multi-channel recording informed by the direction of arrival of the target. The proposed method employs a front-end beamformer steered towards the target speaker to provide an auxiliary signal to a single-channel TSE system. By allowing for time-varying embeddings in the single-channel TSE block, the proposed method fully exploits the correspondence between the front-end beamformer output and the target speech in the microphone signal. Experimental evaluation on simulated multi-channel 2-speaker mixtures, in both anechoic and reverberant conditions, demonstrates the advantage of the proposed method compared to recent single-channel and multi-channel baselines.
Explore related subjects
Keep this discovery
Mohamed Elminshawi, Srikanth Raj Chetupalli, Emanuël A. P. Habets. 2023-03-15. Beamformer-Guided Target Speaker Extraction. https://arxiv.org/abs/2303.08702
Cite the original work for its findings. Save a collection to share your selection of sources.