arXiv · 2609.16855
Differentiable and Severity-invariant Discrete Tokens for Dysarthric Speech Recognition
Abstract
This paper proposes novel differentiable and severity-invariant (DSI) discrete token approaches that are not only tightly integrated with downstream dysarthric speech recognition tasks, but also minimise discrete token diversity across speech impairment severity groups. Experiments conducted on the UASpeech and TORGO corpora suggest that Conformer models trained using the DSI tokens outperform the comparable baseline HuBERT discrete/continuous features by statistically significant WER reductions of 2.22\%/0.78\% absolute (9.14\%/3.41\% relative) and 1.78\%/1.06\% absolute (18.43\%/11.86\% relative) on the two tasks, respectively. After system combination, the lowest WERs of 18.90\% and 6.38\% were obtained on UASpeech and TORGO. Phoneme-specific T-SNE visualizations show that severity-invariant regularization reduces severity-dependent variation by producing greater overlap and less distinct boundaries among severity-group distributions.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Huimeng Wang, Xurong Xie, Mengzhe Geng, Haoning Xu, Jiajun Deng, Youjun Chen, Chengxi Deng, Xunying Liu. 2026-09-15. Differentiable and Severity-invariant Discrete Tokens for Dysarthric Speech Recognition. https://arxiv.org/abs/2609.16855
Cite the original work for its findings. Save a collection to share your selection of sources.