🤖 AI Summary
Direct Preference Optimization (DPO) penalizes all rejected tokens during training, which risks inadvertently suppressing beneficial behaviors and diminishing overall model capabilities. To address this issue of improper token-level credit assignment, this work proposes a gradient alignment-based token-level reweighting method. By quantifying the alignment between each token's gradient and the preference optimization direction, the proposed approach dynamically modulates the strength of negative signals, thereby enabling more precise fine-grained learning. Experimental results demonstrate that this method yields significant average performance improvements across eleven benchmarks and exhibits enhanced robustness to aggressive hyperparameter configurations.
📝 Abstract
Most preference optimization methods, such as Direct Preference Optimization (DPO), apply preference supervision at the response level, although autoregressive language models are optimized token by token. As a result, all tokens in a rejected response contribute to the negative training signal, including tokens that may encode behavior that is useful for the preferred response. We introduce GAW-PO, a gradient-aligned token reweighting method for DPO that estimates, for each rejected token, whether penalizing it would interfere with the preferred update directions. Tokens whose gradients are strongly aligned with the preferred behavior receive a weaker negative contribution, while conflicting tokens retain a stronger penalty. Our method achieves the highest average performance among the evaluated preference-optimization methods, improving by 0.97 points over standard DPO and 0.65 points over the strongest competing baseline across 11 benchmarks spanning mathematics, reasoning, coding, and question answering. We further show that gradient-aligned weighting is substantially more robust to aggressive preference optimization: as the DPO regularization parameter $β$ decreases, standard DPO degrades sharply, whereas GAW-PO continues to improve. These results suggest that accounting for the interaction between rejected-token updates and preferred behavior provides an effective form of token-level credit assignment for preference optimization.