arXiv · 2609.35005
Sub-Model Short-Term Memory Convolutions for Keyword Spotting Systems on Device
Abstract
Keyword Spotting (KWS) is becoming increasingly important as voice-controlled devices grow more widespread. While voice interaction with smartphones and smart TVs is already common, deploying KWS on heavily resource-constrained edge devices such as wearables remains challenging. These systems must meet high accuracy requirements while operating under strict constraints on computational power, memory footprint, and real-time latency. In this work, we present an application of the STMC (Short-Term Memory Convolutions) framework to adapt a modular CNN model for online, LSTM-like inference. Our approach reduces power consumption and redundant computations while maintaining the stability and simplicity of training CNNs. We achieve up to 82% and 46% MCPS reduction compared to equivalently frequent standard CNN execution and vanilla STMC, respectively. The best configuration achieves 93.8% accuracy on the 11-class Google Speech Commands task and 97.1% on the same task with zero-padded data.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Paweł Warlewski, Artur Czeczko, Artur Szumaczuk, Grzegorz Stefański, Szymon Klimaszewski. 2026-09-28. Sub-Model Short-Term Memory Convolutions for Keyword Spotting Systems on Device. https://doi.org/10.21437/interspeech.2026-1343
Cite the original work for its findings. Save a collection to share your selection of sources.