arXiv · 2610.12161
When KL Regularization Misfires in Group Policy Optimization
Abstract
Why does removing reference-policy KL regularization sometimes improve group policy optimization? This motivates studying how reference-policy information should enter group-relative updates. We analyze seven potential failure modes in the interactions between KL and rewards: residual KL updates after reward clipping, after gradient cancellation, and in groups with identical rewards; KL growth with response length and an imbalance in its relative contribution; KL concentration on a small number of tokens; and sampling noise when k1 is incorporated into rewards. We propose Zero-Sum Calibrated Policy Optimization (ZCPO), which uses relative drift measured by conditional KL to calibrate within-group reward coefficients and integrates them into the base surrogate. Mathematical reasoning experiments and ablations support this design's effectiveness in our settings.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Fei Ding. 2026-10-08. When KL Regularization Misfires in Group Policy Optimization. https://arxiv.org/abs/2610.12161
Cite the original work for its findings. Save a collection to share your selection of sources.