arXiv · 2406.09592
On Value Iteration Convergence in Connected MDPs
Abstract
This paper establishes that an MDP with a unique optimal policy and ergodic associated transition matrix ensures the convergence of various versions of the Value Iteration algorithm at a geometric rate that exceeds the discount factor {\gamma} for both discounted and average-reward criteria.
Explore related subjects
Keep this discovery
Arsenii Mustafin, Alex Olshevsky, Ioannis Ch. Paschalidis. 2024-06-13. On Value Iteration Convergence in Connected MDPs. https://arxiv.org/abs/2406.09592
Cite the original work for its findings. Save a collection to share your selection of sources.