arXiv · 1509.01644
Reinforcement Learning with Parameterized Actions
Abstract
We introduce a model-free algorithm for learning in Markov decision processes with parameterized actions-discrete actions with continuous parameters. At each step the agent must select both which action to use and which parameters to use with that action. We introduce the Q-PAMDP algorithm for learning in these domains, show that it converges to a local optimum, and compare it to direct policy search in the goal-scoring and Platform domains.
Explore related subjects
Keep this discovery
Warwick Masson, Pravesh Ranchod, George Konidaris. 2015-09-05. Reinforcement Learning with Parameterized Actions. https://arxiv.org/abs/1509.01644
Cite the original work for its findings. Save a collection to share your selection of sources.