🤖 AI Summary
This study addresses the issue that high-frequency symmetric tokens dilute preference signals in Direct Preference Optimization (DPO), leading to gradient entanglement and degraded optimization efficacy. To mitigate this, we propose Anisotropic DPO, which introduces a vocabulary frequency-based hard masking mechanism coupled with a non-uniform weighting strategy. This approach suppresses the reward contributions of high-frequency shared tokens, thereby amplifying discriminative preference signals. Notably, the method incurs no additional parameters or computational overhead, as it solely adjusts token-level objective weights to alleviate gradient interference. Experimental evaluations on benchmarks such as AlpacaEval demonstrate that Anisotropic DPO significantly outperforms standard DPO, effectively enhancing model robustness against noise. Ultimately, this work provides a zero-cost, plug-and-play solution for preference optimization.
📝 Abstract
Direct Preference Optimization (DPO) aligns language models by optimizing over sequence-level sums of token-wise implicit reward differences. However, we identify a pervasive pathology in this formulation: a disproportionately small subset of high-frequency token types dominates cumulative sequence scores while appearing symmetrically across both preferred and dispreferred responses. Specifically, under canonical Qwen tokenization on Anthropic HH-RLHF, merely 69 token types account for $55.1\%$ of all response tokens and $85.9\%$ of within-pair shared token mass, exhibiting substantially lower preference-side specificity than the remaining vocabulary. This symmetric ubiquity induces gradient entanglement and dilutes the discriminative preference signal propagated through the objective. To resolve this issue, we introduce \emph{Anisotropic DPO} (\textsf{ADPO}) and its canonical realization, \emph{Frequency-Hard DPO}. Using a fixed, label-agnostic vocabulary mask, our method zeroes the implicit reward contribution of high-frequency response tokens while assigning unit weight to informative positions, thereby suppressing gradient interference without modifying preference pairs, discarding context, or introducing learned parameters. Here, \emph{anisotropy} designates non-uniform token-level objective weighting rather than representational geometry. Extensive empirical evaluations on AlpacaEval, MT-Bench, and Arena-Hard demonstrate that Frequency-Hard DPO consistently outperforms standard DPO across Qwen-2.5-7B-Instruct and Llama-3-8B-Instruct, establishing that selectively masking shared high-frequency tokens offers an effective, zero-overhead mechanism for robust preference alignment.