arXiv · 2607.15067
Kernel weighted importance sampling for off-policy evaluation in contextual bandits
Abstract
This article presents a novel estimator for performing off-policy evaluation using only offline data for contextual bandits. The proposed estimator, Kernel-WIS is demonstrated to be asymptotically consistent and to empirically outperform strong baselines (including weighted importance sampling), particularly under behaviour policy miss-specification. The benefit of Kernel-WIS is derived from combining the bounded property of weighted importance sampling with the linearity of vanilla importance sampling.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Joshua Spear, Matthieu Komorowski, Rebecca Pope, Erica E. M. Moodie. 2026-07-16. Kernel weighted importance sampling for off-policy evaluation in contextual bandits. https://arxiv.org/abs/2607.15067
Cite the original work for its findings. Save a collection to share your selection of sources.