Measuring the Assistant's Harmlessness Preferences on the User Turn

📅 2026-09-20
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
研究探讨了通过后训练,模型不仅在助手发言时体现无害偏好,在用户发言时也受影响。此现象表明后训练使模型更广泛地理解用户角色。
📝 Abstract
Post-training turns a general next-token predictor into a chat model with a persistent assistant persona. If that persona is a character the model plays only on its own turns, its preferences should govern what the assistant says, not what the model predicts other speakers will say. We test this boundary and find that it does not hold: a safety-relevant preference of the assistant---for harmless over harmful tasks---shapes the model's predictions even on the user's turn, where the assistant is not the one speaking. We find that this preference is small or near-zero in pretrained base models, that it emerges through post-training, replicated across open-weight model families, grows with scale, and can be moved by narrow finetuning that never touches user turns. We claim that this is evidence that post-training does not merely install a shallow assistant persona, but instead generalises beyond just the local assistant turn, into the model's representation of the user.
Problem

Research questions and friction points this paper is trying to address.

safety-relevant preference
harmless over harmful tasks
post-training
assistant persona
model's predictions
Innovation

Methods, ideas, or system contributions that make the work stand out.

post-training
assistant persona
harmlessness preference
model representation of the user