arXiv · 2606.06357
F3-Tokenizer: Taming Audio Autoencoder Latents for Understanding and Generation
Abstract
Continuous audio autoencoders reconstruct waveforms well but often produce latents with weak structure for understanding, while self-supervised audio encoders capture semantics but are not directly decodable. This mismatch complicates a single audio tokenizer that must support both understanding and generation. We adapt continuous autoencoder latents to this setting with two components: a noise-regularized autoencoder bottleneck and a latent-side representation encoder. The bottleneck uses channel normalization and stochastic perturbation instead of KL-based variational training, yielding scale-controlled continuous latents for reconstruction and autoregressive generation. The representation encoder is trained on frozen autoencoder latents with RQ-MTP and frozen-LLM supervision. The resulting tokenizer provides high-dimensional representations for understanding while preserving normalized continuous latents as generation targets
Explore related subjects
Keep this discovery
Dinghao Zhou, Xingchen Song, Di Wu, Pengyu Cheng, Shengfan Shen, Sixiang Lv. 2026-06-04. F3-Tokenizer: Taming Audio Autoencoder Latents for Understanding and Generation. https://arxiv.org/abs/2606.06357
Cite the original work for its findings. Save a collection to share your selection of sources.