arXiv · 2610.02614
What Is Lost in Post-Training? Default Collapse and the Loss of In-Context Steerability Across Diverse Perspectives
Abstract
AI models serving a heterogeneous population must act on the principles appropriate to each user and context. While post-training has been shown to narrow the views large language models express, prior work has focused on default behavior rather than the ability to adapt to in-context information. We show that post-training also degrades a model's ability to be steered in-context toward perspectives it was not trained to favor. In controlled experiments, we fine-tune models toward one side of cultural-value disagreements and evaluate checkpoints throughout training. The trained side becomes increasingly dominant in ordinary use, while the ability to recognize and faithfully enact the opposing view declines. These findings point to a tension between prioritizing a single set of values and preserving the technical capacity needed to serve diverse stakeholders. Finally, we propose and analyze an alternative objective that maximizes reward subject to a prescribed distribution over expressed perspectives, and present stance-distribution matching as a practical implementation.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Jessica Dierking, Itai Shapira, Niclas Boehmer. 2026-10-02. What Is Lost in Post-Training? Default Collapse and the Loss of In-Context Steerability Across Diverse Perspectives. https://arxiv.org/abs/2610.02614
Cite the original work for its findings. Save a collection to share your selection of sources.