arXiv · 2609.34337
Audio Tokens as a Budgeted Resource: Marginal-Utility Allocation for Scalable Audio Representations
Abstract
Discrete audio tokens are widely used as a representation interface, yet fixed-depth RVQ tokenizers allocate equal capacity to every frame despite varying refinement value. We introduce UniAdapt, which learns marginal utility of RVQ refinements on a frozen codec and allocates them under exact serialized-bit budgets. A rate-independent causal controller predicts acoustic utility, while an optional semantic head supports speech-only utterance-level allocation; measured acoustic and semantic marginal gains on speech have a correlation of 0.42. For causal allocation, a primal-dual allocator selects prefix-valid depths, while an exact guard constrains each sequence prefix to its matched fixed-depth serialized budget. Under utterance-level allocation, UniAdapt reduces Log-STFT distortion by 1.07-4.39 percent across speech, music, and environmental audio without larger budgets. Causally, it improves three of four speech rates with zero violations across 800 utterance-rate evaluations and runs faster than real time. A 20-listener utterance-level MUSHRA study shows a significant 3.52-point speech improvement, with no significant differences on music or environmental audio. These results support separating utility prediction from budget enforcement for scalable, budget-conditioned audio representations.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Mingyu Zhao, Jinchao Zhang, Zhiyong Wu. 2026-09-28. Audio Tokens as a Budgeted Resource: Marginal-Utility Allocation for Scalable Audio Representations. https://arxiv.org/abs/2609.34337
Cite the original work for its findings. Save a collection to share your selection of sources.