What Is Lost in Post-Training? Default Collapse and the Loss of In-Context Steerability Across Diverse Perspectives

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the phenomenon of "default collapse" induced by post-training in large language models, wherein reinforcing specific values compromises context steerability and hinders the model's capacity to adopt opposing perspectives for diverse users. We present the first quantitative analysis of this adverse effect on context steerability, elucidating its underlying mechanisms through controlled fine-tuning experiments and checkpoint evaluations. To mitigate this issue, we propose an alternative objective based on maximizing a predefined perspective distribution, alongside a stance distribution matching scheme. Our findings confirm the inherent tension between value alignment and multi-perspective adaptability, offering an effective technical pathway to restore context steerability while preserving safety alignment.
📝 Abstract
AI models serving a heterogeneous population must act on the principles appropriate to each user and context. While post-training has been shown to narrow the views large language models express, prior work has focused on default behavior rather than the ability to adapt to in-context information. We show that post-training also degrades a model's ability to be steered in-context toward perspectives it was not trained to favor. In controlled experiments, we fine-tune models toward one side of cultural-value disagreements and evaluate checkpoints throughout training. The trained side becomes increasingly dominant in ordinary use, while the ability to recognize and faithfully enact the opposing view declines. These findings point to a tension between prioritizing a single set of values and preserving the technical capacity needed to serve diverse stakeholders. Finally, we propose and analyze an alternative objective that maximizes reward subject to a prescribed distribution over expressed perspectives, and present stance-distribution matching as a practical implementation.
Problem

Research questions and friction points this paper is trying to address.

Post-training
In-context steerability
Default collapse
Large language models
Value alignment
Innovation

Methods, ideas, or system contributions that make the work stand out.

Post-training
In-context steerability
Default collapse
Stance-distribution matching
Large language models
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
J
Jessica Dierking
Hasso Plattner Institute, University of Potsdam
I
Itai Shapira
Harvard University
Niclas Boehmer
Niclas Boehmer
Hasso Plattner Institute
Computational Social ChoiceAI for Social GoodAlgorithmic Game TheoryParameterized Complexity