arXiv · 2610.09448
Beyond Token Revision: Investigating Mask-and-Replace Diffusion for Zero-Shot Text-to-Speech
Abstract
Unlike autoregressive models, discrete diffusion-based models for zero-shot text-to-speech generate speech tokens in parallel and can revisit earlier predictions. Mask-and-replace training extends mask-only training by randomly replacing some tokens, and its gains are commonly attributed to self-correction, the ability to revise previously generated tokens. However, exposure to randomly perturbed context during training may itself improve generation, raising the question of whether these gains require inference-time token revision. To investigate this question, we use DeMaR, which combines mask-and-replace training with confidence-ranked mask-only sampling while preserving the total training corruption probability. Trained from scratch on LibriTTS, DeMaR achieves lower word error rates (WER) than autoregressive and mask-only diffusion baselines using the same speech tokenizer. This advantage persists when each token remains unchanged after first being unmasked. Matched training conditions on two heterogeneous speech tokenizers show that both noisy-context augmentation and replacement supervision improve WER under this restriction. These findings demonstrate training-side benefits of replacement beyond enabling inference-time token revision.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Hounsu Kim, Joonyong Park, Yuki Saito, Satoru Fukayama, Juhan Nam. 2026-10-07. Beyond Token Revision: Investigating Mask-and-Replace Diffusion for Zero-Shot Text-to-Speech. https://arxiv.org/abs/2610.09448
Cite the original work for its findings. Save a collection to share your selection of sources.