arXiv ScienceSearch

arXiv · 2604.00840

Fokker-Planck Analysis and Invariant Laws for a Continuous-Time Stochastic Model of Adam-Type Dynamics

Abstract

We develop a continuous-time model for the long-term dynamics of adaptive stochastic optimization, focusing on bias-corrected Adam-type methods. Starting from a finite-sum setting, we identify a canonical scaling of learning rates, decay parameters, and gradient noise that yields a coupled, time-inhomogeneous stochastic differential equation for the parameters $x_t$, first-moment tracker $z_t$, and second-moment tracker $y_t$. Bias correction persists via explicit time-dependent coefficients, and the dynamics becomes asymptotically time-homogeneous. We analyze the associated Fokker-Planck equation and, under mild regularity and dissipativity assumptions on $f$, prove existence and uniqueness of invariant measures. Noise propagation is governed by $A(x)=\mathrm{Diag}(\nabla f(x))H_f(x)$. Hypoellipticity may fail on $\mathcal D_A\times\mathbb R^m\times(\mathbb R_+)^m$, where \[ \mathcal D_A=\{x\in\mathbb R^m:\exists j,\ e_j^\top A(x)=0\}\subset\{x:\det A(x)=0\}=\mathcal D_A^\dagger, \] and critical points of $f$ lie in $\mathcal D_A$. We show $\mathcal D_A^\dagger\neq\mathbb R^m$ and use this to prove exponential convergence of the Markov semigroup $μ_0P_t$ to a unique invariant measure, uniformly in $μ_0$. The proof uses a Harris-type argument, minorization on Lyapunov sublevel sets, control constructions, and hypoellipticity on $(\mathbb R^m\setminus\mathcal D_A)\times\mathbb R^m\times(\mathbb R_+)^m$. This provides a transparent continuous-time view of Adam-type dynamics.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Kaj Nyström. 2026-04-01. Fokker-Planck Analysis and Invariant Laws for a Continuous-Time Stochastic Model of Adam-Type Dynamics. https://arxiv.org/abs/2604.00840

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Two-layers neural networks for Schr{ö}dinger eigenvalue problems

The aim of this article is to analyze numerical schemes using two-layer neural networks with infinite width for the resolution of high-dimensional Schr{ö}dinger eigenvalue problems with smooth interaction potentials and Neumann boundary condition on the unit cube in any dimension. More precisely, any eigenfunction associated to the lowest eigenvalue of the Schr{ö}dinger operator is a unit L 2 norm minimizer of the associated energy. Using Barron's representation of the solution with a probability measure defined on the set of parameter values and following the approach initially suggested by Bach and Chizat [1], the energy is minimized thanks to a constrained gradient curve dynamic on the 2-Wasserstein space of the set of parameter values defining the neural network. We prove the existence of solutions to this constrained gradient curve. Furthermore, we prove that, if it converges, the represented function is then an eigenfunction of the considered Schr{ö}dinger operator. At least up to our knowledge, this is the first work where this type of analysis is carried out to deal with the minimization of non-convex functionals.

math.AP

Validity of Prandtl Expansion for Steady Compressible Navier-Stokes-Fourier Flows

Assume no-slip boundary conditions for the velocity field and either insulated or Dirichlet boundary conditions for the temperature field in a steady compressible fluid. In the inviscid limit $\v \rightarrow 0$, we develop a mathematical framework for the uniform-in-$\v$ remainder estimate for the linear steady compressible Navier-Stokes-Fourier equations around a Prandtl layer profile with both velocity and thermal layers, which leads to the validity of the Prandtl layer expansion.

math.AP

Long time behaviour of Mean Field Games with fractional diffusion

In this paper we study the long time behaviour of mean field games systems with fractional diffusion, modeling the case that the individual dynamics of the players is driven by independent jump processes and controlled through the drift term, while being confined by an external field in order to guarantee ergodicity. In the case of globally Lipschitz, locally uniformly convex Hamiltonian, and weakly coupled costs satisfying the Lasry-Lions monotonicity condition, we prove that there is a unique solution $(u_T,m_T)$ to the mean field game problem in $(0,T)$ and we show that, if $T$ is sufficiently large, $(u_T,m_T)$ satisfies the so-called turnpike property, namely it is exponentially close to the (unique) stationary ergodic state for any proportionally long intermediate time.

math.AP