arXiv · 2609.29867
Does per-frame early exit pay? A compute-matched study of dynamic depth for on-device speech enhancement
Abstract
Deep learning-based speech enhancement is increasingly deployed on-device in hearing aids, headsets, and earbuds. Most of these devices, however, can only accelerate static int8 graphs, so a depth-varying network must be implemented as several graphs, orchestrated by a policy. In this paper, we supervise every intermediate depth of one causal model, then we fine-tune its output heads to guarantee that deeper outputs are never worse than shallower ones. Using this training protocol, we can derive a family of static models that are more Pareto-efficient than their equivalently-sized counterparts trained from scratch on the same budget. Specifically, we achieve up to 0.11 higher PESQ for equivalent compute, and match the best PESQ at 30% less compute. We then quantize the models to int8 and measure the latency-quality frontier on an STM32N6 microcontroller. On VoiceBank-DEMAND, the dynamic enhancer lies on the same frontier as the static models, rather than trading quality for dynamic execution. Running the policy on the companion Cortex-M55 takes only 26 $μ$s per frame, while splitting the enhancer into separate NPU graphs adds 2.2% latency overhead. The cost of dynamic execution is therefore small.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Clément Laroche, Riccardo Miccini. 2026-09-24. Does per-frame early exit pay? A compute-matched study of dynamic depth for on-device speech enhancement. https://arxiv.org/abs/2609.29867
Cite the original work for its findings. Save a collection to share your selection of sources.