arXiv · 2609.23525
Beyond Encoder Fusion: Multi-View Discrete Token Augmentation for LLM-Based ASR
Abstract
Discrete speech tokens provide a compact interface between speech encoders and large language models for automatic speech recognition, but single-tokenization systems remain sensitive to the chosen encoder. We propose multi-view discrete token augmentation, a simple strategy that augments each training utterance by generating alternative token sequences from fixed SSL encoders, such as HuBERT, WavLM, and MMS-300M. These tokenizations are treated as complementary training views for a shared LLM decoder, exposing it to more diverse discrete speech representations without requiring multi-encoder inference. At test time, the model can operate with a single encoder. On LibriSpeech, the approach consistently improves all encoders over independently trained baselines, with WavLM reaching 3.30% WER on test-clean and 8.13% on test-other. Budget-matched controls show that the gains come from encoder diversity rather than data volume. ROVER over multi-view hypotheses further improves WER to 3.03% and 7.38%.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Paul Moïse Gangbadja, Mickael Rouvier, Fabrice Lefèvre. 2026-09-20. Beyond Encoder Fusion: Multi-View Discrete Token Augmentation for LLM-Based ASR. https://arxiv.org/abs/2609.23525
Cite the original work for its findings. Save a collection to share your selection of sources.