arXiv · 2609.27925
Learning When Not to Listen: Selective Anti-Interference Pretraining for Language Models
Abstract
Language models can over-condition on irrelevant preceding text: predictions already supported by local context may still change when distant, unrelated prefix tokens are perturbed. This interference is especially consequential in long, packed, or distractor-heavy contexts, where useful evidence and irrelevant spans coexist. We propose Selective Prefix Anti-Interference Regularization (SPAR), a pretraining objective for selective anti-interference. SPAR runs the original sequence and a corrupt-prefix input in which only the far prefix is changed, then uses a short-context sufficiency gate and a gated KL objective to stabilize locally supported suffix predictions. The gate operationalizes a model-based estimate of whether the far prefix supplies additional information about the target token. Mechanism analyses show that the gate identifies locally sufficient tokens and sharply reduces prefix sensitivity on gate-selected suffix tokens. In continued training on pretrained base models, SPAR improves RULER across Qwen2.5-0.5B, Qwen2.5-3B, Llama-3.2-1B, Llama-3.1-8B, and GPT2-XL under equal counted training compute; pretraining experiments further show gains on both RULER and NoLiMa. These results show that selective anti-interference is an effective objective-level signal for robust context use.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Jinchang Zhu, Haowei He, Yi Ding, Rong Fu, Nie Xiaojian, Shuangyong Song, Zhongjiang He, Menglin Yang. 2026-08-22. Learning When Not to Listen: Selective Anti-Interference Pretraining for Language Models. https://arxiv.org/abs/2609.27925
Cite the original work for its findings. Save a collection to share your selection of sources.