arXiv · 2404.14442
A Smooth Polynomial Lyapunov Certificate for Convergence of Q-Learning and Its Smooth Variants
Abstract
Classical convergence analyses of Q-learning rely on the $\infty$-norm contraction of Bellman operators, and existing ordinary differential equation (ODE) arguments often use the non-differentiable $\infty$-norm directly. This paper develops a smooth polynomial Lyapunov-function-based stability certificate for convergence of Q-learning by transferring $\infty$-norm contraction to a weighted degree-$2p$ polynomial Lyapunov function induced by a finite $2p$-norm. The framework is conceptual and structural: it avoids non-differentiability, handles preconditioned dynamics arising in Q-learning and its variants, and gives a unified stability argument for standard Q-learning and smooth variants based on log-sum-exp (LSE), mellowmax, and Boltzmann softmax operators. For contractive operators, including the max, LSE, and mellowmax cases, the associated ODEs are globally exponentially stable and, under the stated independent and identically distributed (i.i.d.) sampling model, the stochastic approximation iterates converge almost surely. For the Boltzmann operator, which need not be contractive, the same framework yields convergence to an explicit invariant error set around the optimal Q-function. The resulting theory is not intended as a finite-time bound, but as a clean ODE foundation that unifies and simplifies asymptotic analyses of Q-learning and its smooth variants.
Explore related subjects
Keep this discovery
Donghwan Lee, Hyunjun Na. 2024-04-20. A Smooth Polynomial Lyapunov Certificate for Convergence of Q-Learning and Its Smooth Variants. https://arxiv.org/abs/2404.14442
Cite the original work for its findings. Save a collection to share your selection of sources.