arXiv · 2609.35525
Deep Epistemic Value Functions for Optimistic Exploration
Abstract
Principled exploration in reinforcement learning requires an agent to quantify its epistemic uncertainty and act to resolve it. Uncertainty over the value function provides a natural signal for exploration, yet existing deep approximations remain brittle and perform inconsistently. The central challenge is therefore to scale these ideas robustly. We conduct a systematic empirical study of how epistemic uncertainty is represented, propagated, and optimized in deep epistemic value functions, and uncover distinct failure modes along each of these axes. These findings motivate DEVOTE, a model-free reinforcement learning algorithm that controls how uncertainty generalizes beyond observed data, stabilizes its temporal propagation, and preserves adaptation to the resulting non-stationary exploration objective. Across reward-free exploration and challenging continuous-control tasks, DEVOTE reaches novel states more effectively and achieves higher task return than strong model-free and model-based exploration baselines. These results provide evidence that deep epistemic value functions are a promising path toward scalable, principled exploration.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Leander Diaz-Bone, Marco Bagatella, Jonas Hübotter, Andreas Krause. 2026-09-28. Deep Epistemic Value Functions for Optimistic Exploration. https://arxiv.org/abs/2609.35525
Cite the original work for its findings. Save a collection to share your selection of sources.