arXiv · 2609.40149
Role-Adaptive Policy Optimization for Offline Reinforcement Learning
Abstract
Policy regularization in offline reinforcement learning balances policy improvement against reliance on uncertain value estimates. This balance can differ between selecting actions for execution and supplying actions for critic bootstrapping, yet methods such as TD3+BC couple these roles through a shared policy. We propose Role-Adaptive Policy Optimization (RAPO), which adapts policy-update coefficients according to their roles in value learning and execution. RAPO learns these coefficients by differentiating through candidate policy updates formed using the base algorithm's actor objective. For TD3+BC, RAPO separates bootstrap and execution actors and adapts their coefficients independently: the bootstrap objective penalizes policy-induced changes in target values, while the execution objective evaluates a local policy-improvement surrogate. For IQL, whose value learning is already independent of the execution actor, RAPO preserves the original value updates and adapts only the inverse temperature in advantage-weighted policy extraction. Experiments on D4RL locomotion and AntMaze tasks show improvements over both base algorithms, with larger gains for TD3+BC, whose RAPO instantiation outperforms baselines on average.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Seonvin Cho, Soohyun Choi, Songnam Hong. 2026-09-30. Role-Adaptive Policy Optimization for Offline Reinforcement Learning. https://arxiv.org/abs/2609.40149
Cite the original work for its findings. Save a collection to share your selection of sources.