arXiv · 2609.36486
Optimal Multi-Reward Reinforcement Learning
Abstract
We study an unknown-transition finite-horizon Markov decision process (MDP) with a finite collection of known reward functions $\{r^1, r^2, \ldots, r^M\}$. The goal is to output an $ε$-optimal policy for every reward using online episodic interaction only. Performance is measured by the policy error $V_{0}^{*, m} - V_{0}^{\widehatπ^{m}, m}$ where $m\in [M]$ represents the reward function and $V_{0}^{*, m}=\mathbb{E}_{s_1\sim μ}[V_{1}^{*, m}(s_1)]$. Under this setting, we design a provably efficient algorithm to establish a minimax sample complexity bound of $$ O\left(\frac{SAH^3}{ε^2}\log M \mathrm{polylog}\left(\frac{SAH\log M}{\min\left\{ε, 1\right\}δ}\right)\right)$$ episodes, with no additional burn-in cost. This matches the information-theoretic lower bound up to a factor of $ \mathrm{polylog}(SAH\log M/(\min\left\{ε, 1\right\}δ))$. Our method combines three technical ingredients. First, we adapt MVP to reward-switching learning to construct optimistic value estimates. Second, we use fresh replay samples to conservatively evaluate the candidate policies. Third, gap-based multiplicative weights updates adjust the reward-sampling distribution using the differences between these estimates, converting weighted learning progress into simultaneous guarantees for all rewards.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Zijun Chen, Zihan Zhang. 2026-09-29. Optimal Multi-Reward Reinforcement Learning. https://arxiv.org/abs/2609.36486
Cite the original work for its findings. Save a collection to share your selection of sources.