arXiv · 2004.05298
Detached Error Feedback for Distributed SGD with Random Sparsification
Abstract
The communication bottleneck has been a critical problem in large-scale distributed deep learning. In this work, we study distributed SGD with random block-wise sparsification as the gradient compressor, which is ring-allreduce compatible and highly computation-efficient but leads to inferior performance. To tackle this important issue, we improve the communication-efficient distributed SGD from a novel aspect, that is, the trade-off between the variance and second moment of the gradient. With this motivation, we propose a new detached error feedback (DEF) algorithm, which shows better convergence bound than error feedback for non-convex problems. We also propose DEF-A to accelerate the generalization of DEF at the early stages of the training, which shows better generalization bounds than DEF. Furthermore, we establish the connection between communication-efficient distributed SGD and SGD with iterate averaging (SGD-IA) for the first time. Extensive deep learning experiments show significant empirical improvement of the proposed methods under various settings.
Explore related subjects
Keep this discovery
An Xu, Heng Huang. 2020-04-11. Detached Error Feedback for Distributed SGD with Random Sparsification. https://arxiv.org/abs/2004.05298
Cite the original work for its findings. Save a collection to share your selection of sources.