How LLM Task-Adaptation Reshapes Alignment: A Multi-dimensional Study of Behavioral and Representational Drift

📅 2026-07-10
🏛️ arXiv.org
📈 Citations: 1
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the unclear and insufficiently evaluated impact of large language model task adaptation on multi-dimensional alignment. It presents the first systematic quantification of alignment drift across six dimensions under three post-training paradigms—Supervised Fine-Tuning (SFT), KL-regularized SFT, and Reinforcement Learning with Verifiable Rewards (RLVR)—while correlating behavioral shifts with representational changes. The findings reveal that RLVR induces minimal alignment drift, whereas SFT causes substantial degradation. Although KL regularization partially mitigates this drift, it remains inferior to RLVR. Furthermore, representational drift is shown to lag behind behavioral changes. By establishing the novel perspective that task adaptation inherently constitutes an alignment intervention, this work provides a standardized evaluation paradigm for assessing multi-dimensional alignment in large language models.
📝 Abstract
Post-training is a key mechanism for adapting large language models to downstream tasks. While prior work suggests that task adaptation can alter a model's pre-existing alignment, especially its safety behavior, its broader effects across alignment domains remain poorly understood. We address this gap through a systematic evaluation of representative task-adaptation methods, including supervised fine-tuning (SFT), KL-regularized SFT, and reinforcement learning with verifiable rewards (RLVR) across 15 alignment aspects spanning six key domains: safety, factuality, stance stability, social harm, controllability, and instructability. Our results reveal that post-training does not reshape alignment uniformly. RLVR improves task performance while inducing comparatively small, but non-zero, metric-specific shifts, while SFT leads to substantially larger alignment drift across domains. KL regularization mitigates this effect: stronger reference-model anchoring reduces alignment drift from the baseline, although KL-SFT still falls short of RLVR in preserving alignment. Representation-level analysis further supports this pattern, with shifts in alignment-relevant representations tracking behavioral drift. Together, these results show that task adaptation is not merely a capability-improving step, but an alignment intervention in its own right, motivating multi-dimensional alignment evaluation as a standard component of post-training pipelines.
Problem

Research questions and friction points this paper is trying to address.

LLM task-adaptation
alignment drift
post-training
behavioral drift
representational drift
Innovation

Methods, ideas, or system contributions that make the work stand out.

Task-Adaptation
Alignment Drift
Multi-dimensional Evaluation
Representation Analysis
Post-training