arXiv · 2609.21383
Prediction Dynamics in Depth-Recurrent Language Models
Abstract
Depth-recurrent language models refine predictions through repeated latent updates. Why can intermediate answers agree with the endpoint while their scores continue to change? We derive a sharp margin characterization that decomposes the conservatism of a magnitude bound into common translation, direction relative to the winner, and the pairing of each competitor's update with its score gap. Across Huginn-3.5B and Ouro-1.4B, accounting for update direction and competitor pairing reduces the mean earliest qualifying depth by a further 22.5-34.4% of the total depth beyond translation removal under full answer-text scoring. This retrospective comparison uses completed trajectories. Substantial contributions also occur under label scoring. For shared predictive distributions, we separate common and contrast motion orthogonally and express the common component through candidate-set mass and within-set concentration. Common and contrast energies can attenuate at different rates, allowing a growing preference-change share to coexist with shrinking absolute updates. These findings explain finite-depth answer preservation through the geometry and composition of observed score changes.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Xinyue Luo, Fei Yu. 2026-09-18. Prediction Dynamics in Depth-Recurrent Language Models. https://arxiv.org/abs/2609.21383
Cite the original work for its findings. Save a collection to share your selection of sources.