arXiv · 2606.22790
Scaling Audio Models Efficiently: A Joint Study of Compute Constraints and Optimization Behavior
Abstract
Large automatic speech recognition (ASR) models such as Whisper must be deployed across hardware with widely varying memory and inference-speed constraints. We present a compression framework that jointly parametrizes Whisper deployment along \emph{six} dimensions: model size $x_N$, temporal resolution $x_T$, encoder token stride $x_V$, low-rank adaptation capacity $x_R$, weight precision $x_Q$ and sparsity pattern $x_P$. All axes are jointly optimized against three deployment objectives (word error rate, inference FLOPs, and memory footprint) using a non-dominated sorting genetic evolutionary search (NSGA). Across 50 of the 1,680 candidate configurations evaluated, we measure the marginal effect of each axis on the three objectives and identify compression combinations that dominate naive single-axis scaling, and report a consistent negative result: 1:4 structured sparsity fails to recover acceptable accuracy under any tested recovery budget. We report real, measured memory and accuracy figures for genuinely quantized deployment artifacts, and provide a lookup table mapping deployment scenarios (cloud, server, edge, ultra-constrained) to specific axis configurations with their measured accuracy/memory/compute trade-offs
Explore related subjects
Keep this discovery
Vyom Agarwal, Mokshda Gangrade, Siddharth Pal, Jerry Wu. 2026-06-22. Scaling Audio Models Efficiently: A Joint Study of Compute Constraints and Optimization Behavior. https://arxiv.org/abs/2606.22790
Cite the original work for its findings. Save a collection to share your selection of sources.
Discover connections
Connections use source metadata and explicit phrase matches, not verified experimental comparisons.