🤖 AI Summary
This study addresses the phenomenon of "default collapse" induced by post-training in large language models, wherein reinforcing specific values compromises context steerability and hinders the model's capacity to adopt opposing perspectives for diverse users. We present the first quantitative analysis of this adverse effect on context steerability, elucidating its underlying mechanisms through controlled fine-tuning experiments and checkpoint evaluations. To mitigate this issue, we propose an alternative objective based on maximizing a predefined perspective distribution, alongside a stance distribution matching scheme. Our findings confirm the inherent tension between value alignment and multi-perspective adaptability, offering an effective technical pathway to restore context steerability while preserving safety alignment.
📝 Abstract
AI models serving a heterogeneous population must act on the principles appropriate to each user and context. While post-training has been shown to narrow the views large language models express, prior work has focused on default behavior rather than the ability to adapt to in-context information. We show that post-training also degrades a model's ability to be steered in-context toward perspectives it was not trained to favor. In controlled experiments, we fine-tune models toward one side of cultural-value disagreements and evaluate checkpoints throughout training. The trained side becomes increasingly dominant in ordinary use, while the ability to recognize and faithfully enact the opposing view declines. These findings point to a tension between prioritizing a single set of values and preserving the technical capacity needed to serve diverse stakeholders. Finally, we propose and analyze an alternative objective that maximizes reward subject to a prescribed distribution over expressed perspectives, and present stance-distribution matching as a practical implementation.