arXiv · 2609.33298
Direct Hidden-State Alignment: Mapping and Controlling Preference Expression in LLMs
Abstract
In many settings, post-training need not create the target behavior from scratch: the base model can already produce it, but not reliably. This shifts part of preference alignment from capability acquisition to behavioral expression. We ask how a specified preference is represented in native model computation, what prevents target-supporting computation from reliably dominating generation, and whether this structure can directly guide control. We introduce Residual Competition Maps (RCMs), which map a behavioral preference onto signed causal effects of native residual computation. Across preference domains, RCMs reveal coexisting target-supporting and target-competing effects, input-dependent component roles, and cases where a single native-component intervention reverses the preference outcome. DPO substantially reorganizes these effects and can weaken opposition without guaranteeing its removal. We then propose Direct Hidden-State Alignment (DHSA), which treats inference-time hidden states rather than base-model weights as the direct adaptation space. RCM-guided Causal Activation State Transition (CAST) implements DHSA through local state interventions at a small number of preference-relevant interfaces while freezing the base model. With only 256-16,384 controller parameters, CAST reaches DPO-competitive operating points across three preference domains, can complement DPO-trained models, and can be enabled or removed at inference time.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Fansheng Zhang, Shengran Guo, Zexiao Wang, Liang Yuan, Jiyuan Chen, Ruikun Luo. 2026-09-27. Direct Hidden-State Alignment: Mapping and Controlling Preference Expression in LLMs. https://arxiv.org/abs/2609.33298
Cite the original work for its findings. Save a collection to share your selection of sources.