arXiv · 2609.18989
Function Lives Where Variance Doesn't: Task-Weighted Charts of a Language Model's Computation
Abstract
How many dimensions does a language model's computation actually use? The question is ill-posed until one names a functional. Task-weighted charts make it well-posed: low-dimensional coordinate systems fit against a chosen functional of the representation, under the functional's own metric, turning distillation into plain least squares. Across six models from three families, spanning 70m to 7B parameters, next-token prediction needs 70--90% of the residual stream's width to stay within 5% of intact perplexity, a width consumed by the rare tail of language, and the variance profile predicts none of it: two directions carry 90% of GPT-2's activation variance and almost none of its function. Dimension is per-functional: the model's own uncertainty reads from six coordinates where the full predictive distribution needs hundreds; and it grows with depth. The dissociation is exploitable: when only a few dimensions can be kept, charts trained under the functional's metric preserve the model's predictions better than variance-based or optimal linear compression.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Alexandre Quemy. 2026-08-12. Function Lives Where Variance Doesn't: Task-Weighted Charts of a Language Model's Computation. https://arxiv.org/abs/2609.18989
Cite the original work for its findings. Save a collection to share your selection of sources.