Alignment Inertia: Auditing the Durability of Training Data Influence Through Policy Override Resistance

📅 2026-09-23
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
研究通过提出Override Success Rate和alignment inertia,评估系统提示和微调在改变模型先前行为方面的效果,使用TRAK测试适应信号强度。
📝 Abstract
Platform operators increasingly rely on system prompts and fine-tuning to govern model behavior, yet it remains unclear how reliably these interventions override behavior inherited from prior training. We propose Override Success Rate (OSR) and alignment inertia to measure when operator interventions succeed or fail to change prior behavior. We evaluate zero-shot prompting and LoRA fine-tuning across Llama and Mistral in medical misinformation and hate speech. Alignment inertia persists across both models but varies by model, domain, and policy direction. Notably, in Mistral's restrictive hate-speech condition, LoRA increased inertia by 46.5 percentage points, showing that fine-tuning can reinforce rather than override prior behavior. We also use TRAK to test whether inertia is associated with weaker adaptation signals. TRAK achieves AUC of at least 0.85 in 7 of 8 conditions and outperforms model confidence, TF-IDF similarity, and embedding similarity as a predictor of inertia. These results provide an operator-facing audit of where prior training constrains downstream model governance.
Problem

Research questions and friction points this paper is trying to address.

Override Success Rate
alignment inertia
model behavior
fine-tuning
system prompts
Innovation

Methods, ideas, or system contributions that make the work stand out.

Override Success Rate
alignment inertia
LoRA fine-tuning
TRAK
🔎 Similar Papers
No similar papers found.