Which Preferences to Train On? End-to-End Multi-Objective Alignment with an Adversarial Preference Distribution

📅 2026-10-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the insufficient coverage of hard-to-train regions caused by fixed preference distributions in multi-objective alignment of large language models. To this end, it proposes MAESTRO, a framework that optimizes the Pareto front via end-to-end reinforcement learning over a single conditioned policy guided by adversarial preference distributions. The core innovation lies in introducing a Dirichlet-based adversarial distribution dynamically updated through online mirror descent, enabling training to adaptively focus on preference regions where the current policy underperforms. By integrating minimax optimization with prompt-conditioned policies, MAESTRO achieves superior Pareto fronts across multiple tasks at minimal computational cost, substantially improving alignment performance in previously underrepresented and difficult-to-optimize regions.
📝 Abstract
Aligning large language models (LLMs) with human values is important for safe, efficient, and beneficial AI deployment. However, human values are multifaceted: helpfulness, harmlessness and humor trade off against one another, and different users want different trade-offs. Multi-objective alignment (MOA) addresses this by training a policy that can provide any point of the Pareto front, but existing methods either train one model per preference, interpolate a few separately aligned experts post hoc, or train a single conditioned model without considering which preferences it should be trained on. Since the hard regions of the preference simplex depend on the objectives at hand, existing methods leave them under-trained and do not get the most out of a single model. Therefore, we propose MAESTRO (Multi-objective Alignment via End-to-end STeering and Robust Optimization), which formulates MOA as a minimax problem over preference distributions and trains a single prompt-conditioned policy end-to-end with RL against an adversarial preference distribution: a Dirichlet distribution updated by online mirror descent toward the preferences the current policy serves worst, rather than on a fixed one. On HH-RLHF, BeaverTails and a summarization task, with up to three objectives, MAESTRO attains the best Pareto front on most tasks in a single training run, at the lowest training cost among the compared methods. The largest margins appear in the hard regions that a fixed preference distribution leaves under-trained, confirming that a single prompt-conditioned model is capable of covering the objective trade-offs on its own.
Problem

Research questions and friction points this paper is trying to address.

Multi-Objective Alignment
Large Language Models
Preference Distribution
Pareto Front
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multi-Objective Alignment
Adversarial Preference Distribution
Minimax Optimization
Online Mirror Descent
Pareto Front
🔎 Similar Papers
No similar papers found.