arXiv Science⌕ Search

arXiv · 2610.09203

FlexGrad: a root-free approach to adaptive step sizes

Abstract

We introduce the FlexGrad optimizer, an adaptive gradient method that addresses a key limitation of classical AdaGrad-type methods: their inability to exploit correlations between gradients. By scaling steps using both cumulative gradients and cumulative squared gradients, FlexGrad preserves AdaGrad-like behavior in noisy regimes while allowing substantially larger steps when gradient directions are persistent. The method is fully online, horizon-free, and avoids any auxiliary search procedure during optimization. For convex Lipschitz objectives, we prove that FlexGrad attains the optimal asymptotic convergence rate for nonsmooth convex optimization, with guarantees for both global and coordinate-wise variants. FlexGrad is straightforward to implement and is closely connected to practical gradient-normalized methods. Experiments on a range of convex and nonconvex problems show that FlexGrad performs competitively across a range of optimization settings.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Pablo Barros, Aaron Defazio, Vincent Guigues. 2026-10-06. FlexGrad: a root-free approach to adaptive step sizes. https://arxiv.org/abs/2610.09203

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Constrained portfolio game with heterogeneous agents

We investigate stochastic utility maximization games under relative performance concerns in both finite-agent and infinite-agent (graphon) settings. An incomplete market model is considered where agents with power (CRRA) utility functions trade in a common risk-free bond and individual stocks driven by both common and idiosyncratic noise. The Nash equilibrium for both settings is characterized by forward-backward stochastic differential equations (FBSDEs) with a quadratic growth generator, where the solution of the graphon game leads to a novel form of infinite-dimensional McKean-Vlasov FBSDEs. Under mild conditions, we prove the existence of Nash equilibrium for both the graphon game and the $n$-agent game without common noise. Furthermore, we establish a convergence result showing that, with modest assumptions on the sensitivity matrix, as the number of agents increases, the Nash equilibrium and associated equilibrium value of the finite-agent game converge to those of the graphon game.

math.OC↗

A Model-Based Derivative-Free Optimization Algorithm for Partially Separable Problems

We propose UPOQA, a derivative-free optimization algorithm for partially separable unconstrained problems, leveraging quadratic interpolation and a structured trust-region framework. By decomposing the objective into element functions, UPOQA constructs underdetermined element models and solves subproblems efficiently via a modified projected gradient method. Innovations include an approximate projection operator for structured trust regions, improved management of elemental radii and models, a starting point search mechanism, and support for hybrid black-white-box optimization, etc. Numerical experiments on 85 CUTEst problems demonstrate that \texttt{UPOQA} can significantly reduce the number of function evaluations. To quantify the impact of exploiting partial separability, we introduce the speed-up profile to further evaluate the acceleration effect. Results show that the speed-up of UPOQA over baselines is less significant in low-precision scenarios but becomes more pronounced in high-precision scenarios. Applications to quantum variational problems further validate its practical utility.

math.OC↗

Convergence Analysis of Noisy Distributed Gradient Descent for Non-convex Optimization -- Saddle Point Escape

This paper studies noisy distributed gradient descent (\textbf{NDGD}) for smooth non-convex finite-sum optimization over networks. Random perturbations enable saddle-point escape while preserving distributed implementation and consensus. Under suitable regularity conditions, \textbf{NDGD} converges with high probability to a neighborhood of a common local minimizer. Its convergence complexity is comparable to centralized first-order saddle-point escape methods, reducing exponential dependence on problem dimension to polynomial dependence. Numerical experiments demonstrate improved saddle-point escape over standard \textbf{DGD}.

math.OC↗