arXiv · 2603.01875
KDFlow: A User-Friendly and Efficient Knowledge Distillation Framework for Large Language Models
Abstract
Knowledge distillation (KD) is widely used to compress and post-train large language models (LLMs), yet many existing frameworks execute teacher inference with the same training-oriented backend as student optimization, leading to suboptimal efficiency. In this paper, we propose KDFlow, a novel framework for LLM distillation that features a decoupled architecture and employs SGLang for teacher inference. KDFlow combines SGLang for teacher inference with PyTorch FSDP2 for student optimization, allowing each model to run on a backend tailored to its workload. To enable efficient full-vocabulary distillation in this decoupled architecture, KDFlow transfers the teacher's final hidden states via Ray's object store and recomputes teacher logits on each student worker using a frozen copy of the teacher's output head. Furthermore, our framework supports both off-policy and on-policy distillation and incorporates cross-tokenizer algorithms through highly extensible and user-friendly APIs. Experiments show that KDFlow achieves a 1.44$\times$ to 6.36$\times$ speedup over MS-SWIFT in off-policy distillation and a 1.43$\times$ to 1.75$\times$ speedup over verl in on-policy distillation. KDFlow further scales to 64 GPUs, achieving 3.68$\times$ and 2.52$\times$ strong-scaling speedups in two representative model configurations. The code and documentation are publicly available.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Songming Zhang, Xue Zhang, Tong Zhang, Bojie Hu, Yufeng Chen, Jinan Xu. 2026-03-02. KDFlow: A User-Friendly and Efficient Knowledge Distillation Framework for Large Language Models. https://arxiv.org/abs/2603.01875
Cite the original work for its findings. Save a collection to share your selection of sources.
Discover connections
Connections use source metadata and explicit phrase matches, not verified experimental comparisons.