arXiv · 2609.30996
The Linear Representation Hypothesis for Vision-Language-Action Models
Abstract
The linear representation hypothesis (LRH) has become a standard lens for measuring and intervening on semantic information through the internal representations of large language models (LLMs). A growing body of work has begun extending this perspective to vision-language-action (VLA) models, but the dynamical nature of embodied interaction introduces an additional challenge. Unlike semantic attributes commonly studied in LLMs, such as gender or language, a physical quantity of interest (QoI) in a VLA evolves jointly with the system dynamics: the representation influences the actions selected by the policy, which alter the physical state and, in turn, the next representation. In this paper, we develop a theoretical, signature-based formulation of the LRH for VLA that unifies representations and policies. On the representation side, we establish the existence of representations from which the future evolution of a QoI under a candidate action trajectory can be recovered via linear probing. On the policy side, we introduce a signature generalized linear model for stochastic action chunks. This structure yields a monotonic change in the expected future QoI along linear paths in natural parameter space, enabling linear steering. We construct an explicit oracle representation in a planar control-affine navigation experiment and verify the predicted linear probing and steering mechanisms.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Minseok Jeong, Hyewon Choi, Hiroyasu Tsukamoto, SooJean Han. 2026-09-25. The Linear Representation Hypothesis for Vision-Language-Action Models. https://arxiv.org/abs/2609.30996
Cite the original work for its findings. Save a collection to share your selection of sources.