arXiv · 2608.22595
Sparse Additive Off-Policy Evaluation for Reinforcement Learning with Potentially Limited Number of Trajectories
Abstract
We develop a new framework for flexible, nonlinear, and interpretable off-policy evaluation for infinite-horizon reinforcement learning. To handle large state spaces and support transparent decision-making, we model the Q-function using a nonlinear function class with a sparse additive structure. We derive high-probability finite-sample error bounds for estimating the value function of a target policy and show that the bounds depend only logarithmically on the ambient dimension $d$, thereby alleviating the curse of dimensionality. In contrast to most existing theory for off-policy evaluation, which typically assumes access to many trajectories, our analysis guarantees accurate value estimation when either the number of trajectories or the time horizon is sufficiently large. In addition, we propose a group-sparsity-based feature screening procedure that identifies, with high probability, a reduced feature set containing all relevant covariates. Numerical experiments demonstrate the effectiveness of the proposed approach.
Explore related subjects
Keep this discovery
Tuoyi Zhao, Chengchun Shi, Zhengling Qi, Lan Wang. 2026-08-23. Sparse Additive Off-Policy Evaluation for Reinforcement Learning with Potentially Limited Number of Trajectories. https://arxiv.org/abs/2608.22595
Cite the original work for its findings. Save a collection to share your selection of sources.