CODEBLOCK: Learning to Supervise Code at the Right Granularity
Supervised fine-tuning of code LLMs typically applies uniform cross-entropy loss to all response tokens, implicitly assuming that every token provides equally useful learning signals. Recent token-level selection methods challenge this assumption in natural-language SFT by supervising only high-value tokens. However, such pointwise selection can fragment the syntactic structures and program dependencies of code, leaving supervision scattered across incomplete code units. In experiments on 30K code instruction-response pairs, under the same 10% token budget, supervising complete coding blocks improves average performance by 11.7 points over prior approaches that supervise isolated tokens. Motivated by this observation, we propose CodeBlock, a structure-aware sparse supervision framework that uses complete, parser-aligned coding items as the basic units of supervision. CodeBlock constructs coding items from high-quality code instruction data, estimates their supervision utility using GCE, which is more robust to low-probability tokens, and further adjusts their supervision priority using data-flow reach and bridge signals. During training, the full response is retained as context, while loss is applied only to the selected supervision units. Experiments show that partial supervision over only about 6.9% of response tokens consistently outperforms full-token SFT across all five model settings, suggesting that effective code SFT depends not only on identifying high-value supervision, but also on allocating it at the appropriate structural granularity.