arXiv · 2506.13502
BOW: Training Language Models to Reason Over Plausible Next Words
Abstract
Next-word prediction (NWP) trains language models against a single observed continuation, even though many contexts admit multiple plausible next words. Recent RL-based next-word reasoning methods make this tension explicit: they reward a model for producing a rationale that supports one context-conditioned continuation, which can turn a pre-existing preference into a confident, self-justifying trajectory. We introduce BOW, an RL framework that instead trains models to produce self-contained, neutral, and comprehensive descriptions of the plausible next-word space. The policy generates a next-word reasoning trajectory from the full context, but a frozen scorer computes the core reward from that trajectory alone, without separately receiving the context. BOW-Reg adds a lightweight breadth regularizer to this core reward to discourage premature collapse. On two model backbones, BOW remains competitive with the original models and often outperforms trained baselines across ten general reasoning benchmarks. On benchmarks testing ambiguous references and word meanings, BOW-Reg achieves the highest SharedRef correctness and the lowest HoWN-Simple single-sense collapse on both backbones. Human evaluation further shows that BOW-Reg produces broader next-word reasoning trajectories, while direct next-word-prediction evaluation shows that these trajectories remain predictive.
Explore related subjects
Keep this discovery
Ming Shen, Zhikun Xu, Jacob Dineen, Xiao Ye, Ben Zhou. 2026-08-31. BOW: Training Language Models to Reason Over Plausible Next Words. https://arxiv.org/abs/2506.13502
Cite the original work for its findings. Save a collection to share your selection of sources.
Discover connections
Connections use source metadata and explicit phrase matches, not verified experimental comparisons.