Looking in the Mirror: Introspecting Side-Effect Misalignments Induced by Fine-Tuning

📅 2026-08-04
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the unintended degradation of model alignment capabilities—termed “side-effect misalignment”—that can arise when fine-tuning on non-aligned tasks. To mitigate this issue, the authors propose a novel paradigm called “side-effect introspection” and introduce the first benchmark dataset for evaluating such phenomena. The core innovation is the Delta-Aware Introspective Adapter (DAIA), which leverages a LoRA-based architecture to explicitly model the activation discrepancy between the base and fine-tuned models, thereby enhancing the model’s awareness of undesirable alignment shifts. Experimental results demonstrate that DAIA significantly outperforms existing introspection methods across both unseen fine-tuned models and novel safety categories, confirming its strong generalization and effectiveness in preserving alignment integrity during task adaptation.
📝 Abstract
Fine-tuning enables a source model to acquire desired capabilities and behaviors in a target domain while retaining much of its general-purpose competence. However, this adaptation process can also degrade alignment properties that were present in the source model. Recent work has shown that large language models can be trained using LoRA-based modules known as introspection adapters (IAs) to describe behavioral changes induced by fine-tuning. However, existing studies primarily consider settings in which the model is fine-tuned on datasets explicitly designed to implant a specific behavior and is then asked to explain the implanted behavior. This differs from practical deployment scenarios, where the central concern is often side-effect misalignment: unintended degradation of alignment caused by fine-tuning on tasks that are not obviously related to safety or alignment. To bridge this gap, we formulate a novel problem setting called \emph{side-effect introspection}, in which the target of introspection is not a behavior explicitly implanted through fine-tuning, but rather alignment shifts that emerge as unintended side effects, and we construct a dataset for this setting. Furthermore, to enhance sensitivity to internal model changes, we propose the Delta-Aware Introspection Adapter (DAIA), a novel mechanism designed to explicitly process both base-model activations and activation differences induced by fine-tuning. Our empirical evaluation shows that introspection learning generalizes to unseen fine-tuned models and safety categories, and that DAIA consistently outperforms existing introspection adapters.
Problem

Research questions and friction points this paper is trying to address.

side-effect misalignment
fine-tuning
alignment degradation
introspection
large language models
Innovation

Methods, ideas, or system contributions that make the work stand out.

side-effect introspection
Delta-Aware Introspection Adapter
fine-tuning misalignment
LoRA-based introspection
activation difference
K
Kotaro Yoshida
Institute of Science Tokyo
L
Laura Gomezjurado Gonzalez
Stanford University
Yukinori Yamamoto
Yukinori Yamamoto
Oak Ridge National Laboratory
materials
Y
Yuji Naraki
ZOZO Research
R
Ryotaro Shimizu
ZOZO Research
Wenya Wang
Wenya Wang
Nanyang Technological University
Deep LearningKnowledge ReasoningNatural Language ProcessingSentiment Analysis