Long-Run Dynamics of AdaGrad-Norm: Trajectory Stability and Asymptotic Stationarity
AdaGrad-Norm is widely studied, yet its long-run behavior in smooth nonconvex optimization remains delicate. For the standard current-denominator recursion, the same stochastic gradient determines both the update and its normalizer, rendering the effective stepsize nonpredictable. At the square-root normalization, the quadratic smoothness budget admits only logarithmic control rather than a uniform summability bound. We establish trajectory stability and full-sequence asymptotic stationarity directly for this unmodified recursion. Under global smoothness, non-flatness at infinity, conditional unbiasedness, and an affine conditional second-moment bound, we prove \(\E[\sup_{n\geq1} g(θ_n)]<\infty\) and \(\E[\sup_{n\geq1}\|θ_n\|]<\infty\), without assuming bounded iterates. For stationarity, we use a local conditional relative-moment condition that permits unbounded oracle values and covers fixed-size finite-sum mini-batches and locally nondegenerate additive noise. Together with a weak Sard condition, it yields \(\|\nabla g(θ_n)\|\to0\) almost surely and \(\E\|\nabla g(θ_n)\|^2\to0\). The proof combines a current-sample Lyapunov correction, a stopped-excursion stability argument, and finite-window control for the actual nonpredictable stepsize. We also analyze power normalization with \(q>1/2\), where the quadratic normalization budget is summable, identifying \(q=1/2\) as the boundary of the proof mechanism developed here rather than a universal algorithmic phase transition.