arXiv · 2609.24821
The Answer-Basin Representation Hypothesis: We Are Not Probing or Steering Concepts
Abstract
The Linear Representation Hypothesis associates high-level concepts with directions in language models, but it remains unclear how these concept-related linear structures are organized within the model. We propose the Answer-Basin Representation Hypothesis: the probability measure induced over answers by the model's continuation distribution organizes these linear structures, with its statistics represented along linear directions shared across questions. All continuations yielding the same answer form an answer basin, whose mass is their total probability. These basin masses define the pushforward probability measure over answers. We posit that concept-related linear structure emerges from differences in the answer measure rather than being determined by changes in concept labels. Experiments across models and tasks link concept-consistent effects and their reversals in probing and steering to the alignment between concept labels and the answer measure.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Manjiang Yu, Hongji Li, Zihan Wang, Junwei Chen, Xue Li, Priyanka Singh, Yang Cao, Lijie Hu. 2026-09-21. The Answer-Basin Representation Hypothesis: We Are Not Probing or Steering Concepts. https://arxiv.org/abs/2609.24821
Cite the original work for its findings. Save a collection to share your selection of sources.