arXiv · 2609.21240
Rethinking Music Tokenization: A Semantic Codec toward High-Fidelity LLM Music Generation
Abstract
Discrete audio tokenization has become the critical interface between raw waveforms and autoregressive modeling in recent music generation. As a result, music tokenizers must simultaneously support high-fidelity reconstruction and produce discrete sequences that remain amenable to language modeling. Existing reconstruction-oriented tokenizers often mix musical structure with fine acoustic details, producing high-entropy tokens that are hard to model. In contrast, semantics-guided alternatives are designed for speech and do not fit music well, often hurting reconstruction quality. We address these trade-offs by rethinking music tokenization around a measurable notion of music semantic content grounded in downstream Music Information Retrieval tasks. Guided by this definition, we propose MuSeC, a music semantic codec that factorizes semantic and acoustic content directly from mixed signals without source separation. MuSeC preserves information required for high-fidelity reconstruction while producing more LM-friendly discrete units. Empirically, it improves reconstruction quality and yields more predictable token sequences, providing a practical foundation toward high-fidelity LLM music generation. Demos are available at https://longwaytog0.github.io/MuSeC/.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Huakang Chen, Guobin Ma, Yuepeng Jiang, Dake Guo, Jingbin Hu, Hanke Xie, Wenhao Li, Lingxin Xiong, Jian Zhao, Zhonglin Jiang, Yong Chen, Lei Xie, Pengcheng Zhu. 2026-09-18. Rethinking Music Tokenization: A Semantic Codec toward High-Fidelity LLM Music Generation. https://arxiv.org/abs/2609.21240
Cite the original work for its findings. Save a collection to share your selection of sources.