arXiv · 2504.21323
How to Backdoor Image Knowledge Distillation
Abstract
Knowledge distillation is widely used to transfer behavior from a large teacher model to a smaller student. It is often assumed to be safe when the teacher is clean, because classic backdoor attacks rely on poisoned labels and triggers in supervised training, whereas distillation trains the student to match a teacher's outputs. We show that this assumption can fail when the distillation dataset itself is poisoned. Our attack injects triggered and manipulated images that a clean teacher already predicts as an attacker chosen target label, which causes the student to learn a backdoor even though the teacher remains unaffected. We evaluate this threat across multiple manipulation strategies, including targeted adversarial perturbations and targeted GAN based class transitions, and study how distillation settings influence both accuracy and attack success. We show that the attack remains effective even at a 10% poisoning rates. The results demonstrate that a clean teacher alone is not a sufficient safeguard: poisoned distillation data can produce a strongly backdoored student while maintaining competitive performance on clean images. These findings show that the integrity and provenance of distillation data are part of the security boundary of data intensive KD pipelines, even when the teacher itself is trusted.
Explore related subjects
Keep this discovery
Qian Ma, Chen Wu, Prasenjit Mitra, Sencun Zhu. 2025-04-30. How to Backdoor Image Knowledge Distillation. https://arxiv.org/abs/2504.21323
Cite the original work for its findings. Save a collection to share your selection of sources.
Discover connections
Connections use source metadata and explicit phrase matches, not verified experimental comparisons.