arXiv · 2609.32667
Learning from a Thoughtful Teacher: Adaptive On-Policy Self-Distillation for Mathematical Reasoning
Abstract
On-policy self-distillation (OPSD) trains a question-only student with token-level feedback from a teacher given training-only privileged information (PI). OPSD therefore provides dense, on-policy supervision, and is free of a larger external teacher, but its effectiveness rests on how PI is designed and utilized. Our preliminary diagnostics suggest a significant gap between teacher utility and student learnability, where a small fraction of high-disagreement tokens dominate the distillation signal, and short teacher continuations at these positions further expose more explicit PI leakage than transferable correction cues, indicating a strong intent on injecting PI-conditioned shortcuts. We propose Adaptive On-Policy Self-Distillation (AOPSD), which adapts what information the teacher receives and how strongly its feedback influences learning. AOPSD encodes each solution as a reasoning DAG, orders problems by the student's evolving capability, and reveals only the affordable subgraph and its next frontier as PI. For high-disagreement tokens, AOPSD utilizes short teacher continuations as probes to encourage useful guidance while mitigating PI-conditioned shortcuts among teacher supervisions. On HMMT25, AIME24, AIME25, and BRUMo25, AOPSD achieves 72.5% Pass@8, which is 6.7 percentage points above OPSD and 4.2 above the strongest competing baseline while reducing 15 percentage points of training time at lower cost.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Jiacheng Du, Weiwei Xie, Tianyi Du, Shaoxiong Guo, Qibing Ren, Jiaheng Zhang. 2026-09-26. Learning from a Thoughtful Teacher: Adaptive On-Policy Self-Distillation for Mathematical Reasoning. https://arxiv.org/abs/2609.32667
Cite the original work for its findings. Save a collection to share your selection of sources.