Orthogonal Finetuning for Direct Preference Optimization

๐Ÿ“… 2024-09-23
๐Ÿ›๏ธ arXiv.org
๐Ÿ“ˆ Citations: 3
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
DPO models tend to overfit on non-preferred samples, yielding verbose and low-diversity outputs; existing regularization techniques often compromise alignment performance. This paper proposes Rotated Preference Optimization (RoPO), the first preference optimization method leveraging hyperspherical energy invariance to enforce orthogonal regularizationโ€”not via loss modification, but through a rotation-and-scaling weight update mechanism applied directly to the parameter update trajectory. RoPO fine-tunes only 0.0086% of parameters and preserves the original DPO objective unchanged. It maintains strong alignment capability while significantly mitigating overfitting: +10 points on MT-Bench, +2.8 percentage points on AlpacaEval 2, and an average +6-point improvement in generation diversity. RoPO establishes a new paradigm for lightweight, efficient, and alignment-preserving preference optimization.

Technology Category

Search and Optimization: Learning to SearchMachine Learning: Learning Preferences or RankingsIntelligent Robots: Learning & Optimization for ROB

Application Category

Search and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingUser Modeling, Personalization and Recommendation: Fairness-aware retrieval and rankingResponsible Web: Human-perceived consequences of algorithmic deployment on the web
๐Ÿ“ Abstract
DPO is an effective preference optimization algorithm. However, the DPO-tuned models tend to overfit on the dispreferred samples, manifested as overly long generations lacking diversity. While recent regularization approaches have endeavored to alleviate this issue by modifying the objective function, they achieved that at the cost of alignment performance degradation. In this paper, we innovatively incorporate regularization from the perspective of weight updating to curb alignment overfitting. Through the pilot experiment, we discovered that there exists a positive correlation between overfitting and the hyperspherical energy fluctuation. Hence, we introduce orthogonal finetuning for DPO via a weight-Rotated Preference Optimization (RoPO) method, which merely conducts rotational and magnitude-stretching updates on the weight parameters to maintain the hyperspherical energy invariant, thereby preserving the knowledge encoded in the angle between neurons. Extensive experiments demonstrate that our model aligns perfectly with human preferences while retaining the original expressive capacity using only 0.0086% of the trainable parameters, suggesting an effective regularization against overfitting. Specifically, RoPO outperforms DPO by up to 10 points on MT-Bench and by up to 2.8 points on AlpacaEval 2, while enhancing the generation diversity by an average of 6 points.
Problem

Research questions and friction points this paper is trying to address.

Reducing overfitting in DPO-tuned models on dispreferred samples
Maintaining alignment performance while preventing diversity loss
Preserving hyperspherical energy to retain original model knowledge
Innovation

Methods, ideas, or system contributions that make the work stand out.

Orthogonal finetuning via weight-Rotated Preference Optimization
Maintains hyperspherical energy invariant during updates
Conducts rotational and magnitude-stretching weight updates
๐Ÿ”Ž Similar Papers
No similar papers found.
๐Ÿ’ผ Related Jobs
No related jobs found.
Chinese Academy of Sciences | University of Chinese Academy of Sciences | Baidu Inc.
Chenxu Yang
Chenxu Yang
Institute of Information Engineering, Chinese Academy of Sciences
NLPDialogue Generation
R
Ruipeng Jia
Baidu Inc., Beijing, China
N
Naibin Gu
Institute of Information Engineering, Chinese Academy of Sciences, Beijing, China; School of Cyber Security, University of Chinese Academy of Sciences, Beijing, China
Z
Zheng Lin
Institute of Information Engineering, Chinese Academy of Sciences, Beijing, China; School of Cyber Security, University of Chinese Academy of Sciences, Beijing, China
S
Siyuan Chen
Institute of Information Engineering, Chinese Academy of Sciences, Beijing, China; School of Cyber Security, University of Chinese Academy of Sciences, Beijing, China
C
Chao Pang
Baidu Inc., Beijing, China
W
Weichong Yin
Baidu Inc., Beijing, China
Y
Yu Sun
Baidu Inc., Beijing, China
H
Hua Wu
Baidu Inc., Beijing, China
Weiping Wang
Weiping Wang
School of Information Science and Engineering, Central South University
Computer NetworkNetwork Security