Accelerating distributed stochastic gradient descent with adaptive periodic parameter averaging: poster
Abstract
Communication overhead is a well-known performance bottleneck in distributed Stochastic Gradient Descent (SGD), which is a popular algorithm to perform optimization in large-scale machine learning tasks. In this work, we propose a practical and effective technique, named Adaptive Periodic Parameter Averaging, to reduce the communication overhead of distributed SGD, without impairing its convergence property.
DOI 10.1145/3293883.3299818