arXiv · 2610.06675
Improved Convergence of Large Stepsize Gradient Descent for Logistic Regression
Abstract
We study gradient descent (GD) with a large constant stepsize for logistic regression on linearly separable data. Existing analysis shows an accelerated rate of $\widetilde{O}(1/\sqrtε)$ to reach loss $ε$ with an aggressive stepsize, although the loss may initially oscillate. Tighter control of the oscillatory dynamics has been available only for two-dimensional data. We prove a substantially faster rate in arbitrary dimension: GD with a large stepsize $η=1/ε$ reaches loss $ε$ within $O(\ln^{p}(1/ε))$ steps, where $p$ depends only on the margin and the rank of the data. Our proof improves the bound on the transition time of GD from the oscillatory to the stable phase, after which the loss decreases monotonically. We split the oscillatory phase into recursively nested intervals. The margin and the rank bound the nesting depth, and a counting argument bounds the number of intervals at each depth, together yielding the polylogarithmic step complexity.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Xiaochuan Gong, Ang Li. 2026-10-05. Improved Convergence of Large Stepsize Gradient Descent for Logistic Regression. https://arxiv.org/abs/2610.06675
Cite the original work for its findings. Save a collection to share your selection of sources.