Restart and Adaptive Acceleration in Stochastic Gradient Methods
We study restart schemes in stochastic optimization problems for non-smooth and weakly convex functions that satisfy a Kurdyka-Łojasiewicz (KŁ) inequality. Using restarts allows us to leverage KŁ inequalities to achieve optimal convergence rates, with acceleration depending explicitly on the KŁ exponent. Furthermore, for stochastic gradient descent (SGD), optimal restart schedules correspond to piecewise constant (step decay) step sizes. While regularity constants such as the KŁ exponent are typically unknown in practice, we prove that restart schemes are robust to significant misspecification of these constants, hence nearly adaptive. We detail numerical experiments on a number of problems where the KŁ exponent is controlled.