arXiv · 1802.01518
Guided Policy Exploration for Markov Decision Processes using an Uncertainty-Based Value-of-Information Criterion
Abstract
Reinforcement learning in environments with many action-state pairs is challenging. At issue is the number of episodes needed to thoroughly search the policy space. Most conventional heuristics address this search problem in a stochastic manner. This can leave large portions of the policy space unvisited during the early training stages. In this paper, we propose an uncertainty-based, information-theoretic approach for performing guided stochastic searches that more effectively cover the policy space. Our approach is based on the value of information, a criterion that provides the optimal trade-off between expected costs and the granularity of the search process. The value of information yields a stochastic routine for choosing actions during learning that can explore the policy space in a coarse to fine manner. We augment this criterion with a state-transition uncertainty factor, which guides the search process into previously unexplored regions of the policy space.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Isaac J. Sledge, Matthew S. Emigh, Jose C. Principe. 2018-02-05. Guided Policy Exploration for Markov Decision Processes using an Uncertainty-Based Value-of-Information Criterion. https://doi.org/10.1109/tnnls.2018.2812709
Cite the original work for its findings. Save a collection to share your selection of sources.