arXiv · 2609.13911
DualSpecSE: A Dual-Path Speech Enhancement Network Integrating Mel and Complex Spectrograms
Abstract
In this paper, we propose DualSpecSE, a speech enhancement framework that jointly models Mel-spectrogram and complex spectrogram in a dual-path architecture for improved ASR performance and higher-quality speech reconstruction. The Mel branch learns coarse-grained acoustic representations and produces enhanced Mel-spectrograms for direct ASR usage, while the complex branch refines fine-grained spectral details for high-fidelity waveform reconstruction. Built upon the cross-band and narrow-band blocks from CleanMel, DualSpecSE introduces an interaction module and a fusion module to enable effective information exchange between the two branches. The model simultaneously outputs enhanced Mel and complex spectrogram without requiring a pretrained vocoder. Experimental results demonstrate consistent improvements in speech fidelity, perceptual quality, and ASR performance. Codes and audio samples are available.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Xingchen Li, Ziqian Wang, Zikai Liu, Yike Zhu, Zihan Zhang, Longshuai Xiao, Lei Xie. 2026-09-12. DualSpecSE: A Dual-Path Speech Enhancement Network Integrating Mel and Complex Spectrograms. https://arxiv.org/abs/2609.13911
Cite the original work for its findings. Save a collection to share your selection of sources.