arXiv · 2606.29239
QuantGuard: Learnable Rounding for Repairing Quantization-Conditioned Backdoors in LLMs
Abstract
Model quantization is a key technique for reducing storage and inference costs in large language model deployment. However, recent studies show that the discretization and rounding errors introduced by quantization can be exploited by adversaries to construct quantization-conditioned backdoor (QCB) attacks. Under such attacks, malicious behavior remains dormant at full precision and activates only after quantization, thereby bypassing conventional security auditing and detection. To address this threat, we propose QuantGuard, a proactive pre-quantization defense that learns safe rounding adjustments through differentiable optimization. Our method introduces differentiable rounding control variables and combines error-guided rounding reversal constraints, output-distribution consistency, and weight-distance regularization to regulate critical rounding behaviors. Crucially, QuantGuard utilizes only a small calibration dataset and does not modify existing quantization algorithms. This design disrupts the alignment between attacker-crafted weight patterns and quantization boundaries, suppressing post-quantization backdoor activation while preserving model functionality and performance. We conduct systematic experiments on six mainstream LLMs (including the LLaMA-3 and Qwen2.5-Coder) using three quantization precisions (INT8, FP4, and NF4) across three representative scenarios: vulnerable code generation, content injection, and over-refusal. The results show that QuantGuard consistently mitigates QCB attacks, reducing the attack success rate to a level comparable to the clean model while largely preserving general capability. With low computational overhead, QuantGuard provides a practical defense for secure quantized LLM deployment.
Explore related subjects
Keep this discovery
Aoying Zheng, Anqi Du, Zizhuang Deng, Yuxuan Chen, Shanqing Guo, Kening Zheng. 2026-06-28. QuantGuard: Learnable Rounding for Repairing Quantization-Conditioned Backdoors in LLMs. https://arxiv.org/abs/2606.29239
Cite the original work for its findings. Save a collection to share your selection of sources.