arXiv · 2608.27760
Informational Antilocality and the Locality Bias in LLMs
Abstract
We consider the ability of transformer-based language models (LLMs) to learn what we call k-antilocal languages, i.e., languages that have no mutual information across any span of $k$ contiguous symbols. We construct such languages with increasing $k$, finding that LLMs trained on them achieve comparable cross-entropy loss regardless of antilocality, but converge more slowly on more antilocal languages. Our findings support the idea that non-local dependencies are more difficult to learn, but the evidence for this bias comes from learning speed rather than learning success.
Explore related subjects
Keep this discovery
Andrew McInnerney, Shane Storks, Steven Abney, Richard L. Lewis. 2026-08-27. Informational Antilocality and the Locality Bias in LLMs. https://arxiv.org/abs/2608.27760
Cite the original work for its findings. Save a collection to share your selection of sources.
Discover connections
Connections use source metadata and explicit phrase matches, not verified experimental comparisons.