arXiv · 2601.19299
Continuous-time q-learning for Markov regime switching system under Tsallis entropy
Abstract
This paper studies continuous-time q-learning (the continuous-time counterpart of Q-learning) for a Markov regime-switching system under Tsallis entropy regularization. The Tsallis entropy regularization yields an optimal policy distribution that may not necessarily be a Gibbs measure, thereby complicating algorithm design. Furthermore, to address the limited universality of current continuous-time regime-switching reinforcement learning algorithms (often restricted to the exploratory mean-variance framework), this study focuses on continuous-time q-learning for Markov regime-switching systems based on Tsallis entropy, aiming for a more universally applicable continuous-time reinforcement learning method. We establish the martingale characterization of the q-function under Tsallis entropy for continuous-time Markov regime-switching systems. We further design two q-learning algorithms that differ based on whether the Lagrange multiplier can be explicitly derived. We apply these algorithms to the continuous-time exploratory mean-variance portfolio optimization problem in a regime-switching market. Numerical experiments demonstrate the satisfactory performance of our q-learning algorithms.
Explore related subjects
Keep this discovery
Minghui Zhang, Xun Li, Jie Xiong, Xin Zhang. 2026-01-27. Continuous-time q-learning for Markov regime switching system under Tsallis entropy. https://arxiv.org/abs/2601.19299
Cite the original work for its findings. Save a collection to share your selection of sources.