ORPG: Reconciling Multiple Reward Objectives through Objective-wise Policy Gradients

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of aligning learning signals with objectives in multi-reward policy optimization by proposing the ORPG method. Specifically, ORPG constructs independent clipping objectives for each reward and formulates the policy update as the unique solution to a spherical directional compromise. When gradients are compatible, it employs cosine-dependent interpolation for fusion; when conflicts arise, it performs gradient projection according to predefined priorities, thereby achieving synergistic updates within a single policy. Experimental results demonstrate that ORPG significantly outperforms baseline methods in both helpfulness–safety alignment and mathematical reasoning tasks involving correctness–cost trade-offs, effectively improving accuracy while reducing response length.
📝 Abstract
Multi-reward policy optimization requires a joint update that reflects both the learning signals and the intended relationships among objectives. We introduce Objective-wise Reconciled Policy Gradient (ORPG), which constructs a separate clipped policy objective for each reward and reconciles the resulting gradients into one policy update. For compatible gradients, a cosine-dependent interpolation coordinates their contributions through a partially normalized reference while preserving the norm of their sum. We characterize this update as the unique solution of a spherical directional compromise. For conflicting gradients, projection follows the task's priorities. We evaluate the same compatible rule in helpfulness--safety alignment and correctness--cost optimization for mathematical reasoning. ORPG substantially improves average Useful and Harmless scores over the strongest external baseline on each axis. In mathematics, it achieves the highest average full-budget accuracy and three-budget hypervolume among the compared methods, with more accurate and shorter responses than the initial policy. Component comparisons and training dynamics show the larger contribution of compatible coordination and a complementary benefit from conflict handling. These results support gradient reconciliation for objectives with equal standing and for objectives with an explicit priority.
Problem

Research questions and friction points this paper is trying to address.

multi-reward policy optimization
gradient reconciliation
multi-objective alignment
conflicting gradients
Innovation

Methods, ideas, or system contributions that make the work stand out.

Objective-wise Policy Gradients
Gradient Reconciliation
Multi-reward Optimization
Cosine-dependent Interpolation
Conflict Resolution
🔎 Similar Papers
S
Shicheng Fang
Fudan University, Shanghai Innovation Institute
Y
Yiwen Zhao
Fudan University, Shanghai Innovation Institute
W
Wenbo Tian
Fudan University, Shanghai Innovation Institute
J
Jiahao Lu
Fudan University, Shanghai Innovation Institute
Y
Yining Zheng
Fudan University, Shanghai Innovation Institute
Yuxin Wang
Yuxin Wang
Fudan University
X
Xipeng Qiu
Fudan University, Shanghai Innovation Institute