arXiv · 2309.02855
Bandwidth-efficient Inference for Neural Image Compression
Abstract
With neural networks growing deeper and feature maps growing larger, limited communication bandwidth with external memory (or DRAM) and power constraints become a bottleneck in implementing network inference on mobile and edge devices. In this paper, we propose an end-to-end differentiable bandwidth efficient neural inference method with the activation compressed by neural data compression method. Specifically, we propose a transform-quantization-entropy coding pipeline for activation compression with symmetric exponential Golomb coding and a data-dependent Gaussian entropy model for arithmetic coding. Optimized with existing model quantization methods, low-level task of image compression can achieve up to 19x bandwidth reduction with 6.21x energy saving.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Shanzhi Yin, Tongda Xu, Yongsheng Liang, Yuanyuan Wang, Yanghao Li, Yan Wang, Jingjing Liu. 2023-09-06. Bandwidth-efficient Inference for Neural Image Compression. https://arxiv.org/abs/2309.02855
Cite the original work for its findings. Save a collection to share your selection of sources.