arXiv · 1910.07479
Conditional Importance Sampling for Off-Policy Learning
Abstract
The principal contribution of this paper is a conceptual framework for off-policy reinforcement learning, based on conditional expectations of importance sampling ratios. This framework yields new perspectives and understanding of existing off-policy algorithms, and reveals a broad space of unexplored algorithms. We theoretically analyse this space, and concretely investigate several algorithms that arise from this framework.
Explore related subjects
Keep this discovery
Mark Rowland, Anna Harutyunyan, Hado van Hasselt, Diana Borsa, Tom Schaul, Rémi Munos, Will Dabney. 2019-10-16. Conditional Importance Sampling for Off-Policy Learning. https://arxiv.org/abs/1910.07479
Cite the original work for its findings. Save a collection to share your selection of sources.