arXiv · 2609.39701
Values as Style: Disentangling Values from Semantics with One-Way Mixing for Low-Damage LLM Steering
Abstract
Value steering should change an LLM's normative priorities while preserving the scenario, facts, and task constraints underlying its answer. Conventional activation edits often change both. We introduce an editable semantic-value interface on frozen residual states, with a one-way semantic-to-value pathway that grounds value recognition in context. Stop-gradient blocks feedback through this pathway; swap consistency, topic de-confounding, and decorrelation encourage selective codes. At inference, editing the value code produces a residual delta while holding the semantic code fixed. On two instruction-tuned backbones, this interface improves semantic preservation and reduces benign refusals at comparable value alignment. A matched mixing-by-gating ablation separates representation learning from selective edit activation, and dimension-matched probes establish improved code selectivity. Against validation-selected prompting on LLaMA-3.1-8B, the method achieves comparable alignment (0.750 vs. 0.748), higher BERTScore (0.938 vs. 0.923), and fewer contradictions (5.1% vs. 7.6%). Human ratings and cross-taxonomy controls provide complementary evidence for low-damage value steering.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Jiale Dai, Hongcan Deng, Liuxian Ma, Xiaoke Niu, Guojie Song. 2026-09-30. Values as Style: Disentangling Values from Semantics with One-Way Mixing for Low-Damage LLM Steering. https://arxiv.org/abs/2609.39701
Cite the original work for its findings. Save a collection to share your selection of sources.