Robust Nash Alignment under Preference Uncertainty

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the insufficient robustness of preference-based large language model (LLM) alignment methods under uncertainties such as noise, heterogeneity, and distribution shift. To this end, it proposes a game-theoretic optimization framework that formulates a four-player proxy game model, transforming robust alignment into a worst-case win-rate optimization problem. A single-loop optimistic mirror descent-ascent algorithm is designed to solve this formulation efficiently. Theoretically, we establish both the convergence of the proposed algorithm and the existence of an optimal robust strategy. Empirically, experiments demonstrate that our approach significantly outperforms existing baselines in complex preference environments, providing reliable robustness guarantees for LLM alignment.
📝 Abstract
Preference-based alignment methods typically optimize against a single preference model, and can therefore be brittle when pairwise preferences are uncertain: noisy, heterogeneous, or shift after deployment. To address these issues, we propose Robust Nash Alignment, a game-theoretic framework for alignment to uncertain pairwise preferences. Our formulation has a major learner seeking a policy with a large worst-case win rate against both an adversarial competitor and any preference kernel lying in an ambiguity set around a nominal preference. When the ambiguity set captures the uncertainty in preferences, the resulting robust objective of the game directly yields a certified lower bound on worst-case performance. However, we note this problem is computationally challenging to optimize, and to address this, we introduce a four-player primal-dual proxy game involving the leader policy, follower policy, adversarial kernel, and dual variable, and develop a single-loop optimistic mirror descent-ascent algorithm for it. We show that the proxy always lower-bounds the truncated hard-constrained objective, quantify the proxy-to-hard gap, and characterize an exactness condition under which the proxy recovers the robust objective. We then prove an \(\mathcal{O}(1/\sqrt{T})\) average-iteration convergence for the proxy-game duality gap, which implies a near-optimal robust policy for the original robust objective. Experiments on controlled tabular games and LLM alignment with uncertain preference further validate the convergence theory and show improved performance over nominal baselines.
Problem

Research questions and friction points this paper is trying to address.

Preference Uncertainty
Robust Alignment
Nash Equilibrium
Pairwise Preferences
LLM Alignment
Innovation

Methods, ideas, or system contributions that make the work stand out.

Robust Nash Alignment
Preference Uncertainty
Primal-Dual Proxy Game
Optimistic Mirror Descent-Ascent
LLM Alignment