arXiv · 2609.29464
TS-OPD: Reconciling ASR and QA in Speech Language Models via Task-Specific On-Policy Distillation
Abstract
Speech Language Models (SLMs) inherit strong instruction-following capabilities from pretrained language models, yet ASR specialization can substantially degrade them. To address this ASR--QA trade-off, we propose Task-Specific On-Policy Distillation (TS-OPD), which leverages models before and after ASR specialization as complementary QA and ASR teachers. The student generates separate task-conditioned trajectories for ASR and QA, each supervised only by its corresponding teacher, thereby reducing direct competition between the two supervision signals. Experiments on basic ASR, contextual ASR, and QA demonstrate that TS-OPD improves recognition while preserving QA capability. Moreover, TS-OPD remains robust across different balancing coefficients and continues to benefit from increased distillation data.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Yujie Guo, Hongjie Chen, Jian Kang, Jie Li, Yongxiang Li, Yong Qin. 2026-09-24. TS-OPD: Reconciling ASR and QA in Speech Language Models via Task-Specific On-Policy Distillation. https://arxiv.org/abs/2609.29464
Cite the original work for its findings. Save a collection to share your selection of sources.