arXiv · 2304.08048
When do discounted-optimal policies also optimize the gain?
Abstract
In this technical note, we establish an upper-bound on the threshold on the discount factor starting from which all discounted-optimal deterministic policies are gain-optimal, that we prove to be tight on an example. To address computability issues of that theoretical threshold, we provide a weaker bound which is tractable on ergodic MDPs in polynomial time.
Explore related subjects
Keep this discovery
Victor Boone. 2023-04-17. When do discounted-optimal policies also optimize the gain?. https://arxiv.org/abs/2304.08048
Cite the original work for its findings. Save a collection to share your selection of sources.