arXiv · 2609.32674
Expected Reasoning-Step Return Unifies On-Policy Learning from Rewards and Teachers
Abstract
On-policy reasoning models can learn from task rewards or teacher signals, but these sources differ in form and can favor conflicting updates, leaving unclear which should guide a given reasoning action. We introduce \textbf{Expected Reasoning-Step Return (ERSR)}, which treats semantic reasoning steps as macro-actions and uses Monte Carlo student-policy rollouts to estimate the expected final task reward of student-generated and teacher-proposed actions in a common return space for step-level comparison. ERSR analysis reveals an outcome-dependent asymmetry: student actions are more beneficial than teacher replacements on successful trajectories, whereas teacher replacements become more beneficial on failed trajectories. We further show that student answer-probe gains track student-step ERSR utility and distinguish beneficial from harmful reasoning steps. Based on these findings, we propose \textbf{Return-Referenced On-Policy Learning (R$^2$OPL)}, which reinforces student reasoning on successful trajectories and distills teacher signals on failed ones, while using group success rate for difficulty scaling and student-probe gains for step-level modulation. Experiments across reasoning benchmarks and teacher--student configurations show that R$^2$OPL consistently outperforms strong baselines. ERSR training dynamics further show that R$^2$OPL jointly exploits substantial utility from both reward- and teacher-side signals, whereas existing hybrids often leave substantial residual utility in one branch.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Qiangqiang He, Jin Li. 2026-09-26. Expected Reasoning-Step Return Unifies On-Policy Learning from Rewards and Teachers. https://arxiv.org/abs/2609.32674
Cite the original work for its findings. Save a collection to share your selection of sources.