arXiv · 2607.26456
Model-Free Q-Learning for Infinite-Horizon Stochastic Linear Quadratic Problems with Regime Switching
Abstract
This paper addresses infinite-horizon continuous-time stochastic linear quadratic optimal control problems with regime switching. We propose a paradigm shift from model-based design by adopting an adaptive dynamic programming approach, specifically developing on-policy and off-policy Q-learning algorithms that learn the optimal controller solely from online state trajectory data. The theoretical core of our work consists of a complete proof of the equivalence between the on- and off-policy architectures, alongside a rigorous analysis establishing the stability of the closed-loop system and the convergence of the algorithms to the optimal solution. For computational tractability, we implement these algorithms using vectorization and Kronecker product algebra. The theoretical results are corroborated by numerical case studies that clearly demonstrate the operational effectiveness and practical feasibility of the proposed model-free control strategy.
Explore related subjects
Keep this discovery
Xinyue Zhang, Na Li, Xun Li, Zuo Quan Xu. 2026-07-29. Model-Free Q-Learning for Infinite-Horizon Stochastic Linear Quadratic Problems with Regime Switching. https://arxiv.org/abs/2607.26456
Cite the original work for its findings. Save a collection to share your selection of sources.