arXiv · 2105.04646
Deeply-Debiased Off-Policy Interval Estimation
Abstract
Off-policy evaluation learns a target policy's value with a historical dataset generated by a different behavior policy. In addition to a point estimate, many applications would benefit significantly from having a confidence interval (CI) that quantifies the uncertainty of the point estimate. In this paper, we propose a novel deeply-debiasing procedure to construct an efficient, robust, and flexible CI on a target policy's value. Our method is justified by theoretical results and numerical experiments. A Python implementation of the proposed procedure is available at https://github.com/RunzheStat/D2OPE.
Explore related subjects
Keep this discovery
Chengchun Shi, Runzhe Wan, Victor Chernozhukov, Rui Song. 2021-05-10. Deeply-Debiased Off-Policy Interval Estimation. https://arxiv.org/abs/2105.04646
Cite the original work for its findings. Save a collection to share your selection of sources.