arXiv · 2106.15594
Limited depth bandit-based strategy for Monte Carlo planning in continuous action spaces
Abstract
This paper addresses the problem of optimal control using search trees. We start by considering multi-armed bandit problems with continuous action spaces and propose LD-HOO, a limited depth variant of the hierarchical optimistic optimization (HOO) algorithm. We provide a regret analysis for LD-HOO and show that, asymptotically, our algorithm exhibits the same cumulative regret as the original HOO while being faster and more memory efficient. We then propose a Monte Carlo tree search algorithm based on LD-HOO for optimal control problems and illustrate the resulting approach's application in several optimal control problems.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Ricardo Quinteiro, Francisco S. Melo, Pedro A. Santos. 2021-06-29. Limited depth bandit-based strategy for Monte Carlo planning in continuous action spaces. https://arxiv.org/abs/2106.15594
Cite the original work for its findings. Save a collection to share your selection of sources.