arXiv · 2609.37500
REVO: Rollout-Efficient Off-Policy Distillation via Variance-Guided Reuse
Abstract
On-policy distillation (OPD) trains language models using dense token-level teacher supervision on student-generated trajectories. However, its reliance on frequently refreshed student rollouts often incurs substantial generation cost. We introduce REVO, an off-policy distillation framework that improves rollout efficiency by reusing each student rollout for multi-step learner updates. REVO addresses prefix-level and current-token policy mismatch through stabilized prefix weighting and one-step resampling from the current student, which enables repeated updates without regenerating full trajectories. To prioritize informative token positions within reused rollouts, REVO uses the variance of the student-teacher log-probability ratio to quantify the remaining token-level learning signal and guide repeated optimization. Across multiple student-teacher scales, REVO with only 50 rollout iterations matches or exceeds OPD baselines trained for 200 iterations on both in-domain and cross-domain reasoning benchmarks.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Yuxiao Yang, Shangzhe Li, Tianrun Yu, Kaixiang Zhao, Taylor W. Killian, Weitong Zhang. 2026-09-28. REVO: Rollout-Efficient Off-Policy Distillation via Variance-Guided Reuse. https://arxiv.org/abs/2609.37500
Cite the original work for its findings. Save a collection to share your selection of sources.