Groupwise Distortion Guarantees for Preference-Based Alignment

📅 2026-10-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the theoretical gap between social welfare maximization and pairwise comparisons in preference alignment by proposing the GLHF algorithm. This method introduces, for the first time, a group-aware strategy grounded in the Bradley-Terry model, integrating reinforcement learning with Nash equilibrium theory to learn group-conditioned policies that optimize welfare distribution across diverse user populations. The approach maintains sample efficiency while approximating optimal bounds on group distortion. Experimental results demonstrate that GLHF significantly reduces distortion across evaluated groups, outperforming both NLHF and ungrouped baseline methods.
📝 Abstract
Preference-based alignment methods such as reinforcement learning from human feedback (RLHF) and Nash learning from human feedback (NLHF) aggregate pairwise preferences to learn an LLM policy, but a natural goal is maximizing social welfare (average cardinal utility), which comparisons alone need not identify. G\"olz, Haghtalab, and Yang (GHY) measure the gap by distortion: the worst-case ratio between the welfare of the best fixed lottery (distribution over responses) and of the learned lottery. They show NLHF is optimal when every user receives the same lottery. Account-based LLMs, however, have information about their users and can serve different lotteries to different people. We give an efficient algorithm, GLHF, that learns a single group-conditioned policy from one comparison per user. Under individual Bradley--Terry comparisons, GLHF asymptotically matches GHY's optimal population distortion bound simultaneously on every group in a prespecified, possibly overlapping collection, with sample complexity growing logarithmically in the number of groups and inversely with the smallest group mass. A sharper guarantee for groups with similar preferences approaches distortion of one when members share a feasible favorite response. In experiments using human coffee ratings and synthetic LLM-generated ratings, GLHF lowers distortion in every evaluated group and substantially reduces worst-group distortion relative to NLHF and other group-agnostic baselines.
Problem

Research questions and friction points this paper is trying to address.

preference-based alignment
social welfare
distortion
groupwise guarantees
large language models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Groupwise Distortion
Preference-Based Alignment
GLHF Algorithm
Social Welfare
Group-Conditioned Policy
🔎 Similar Papers
2024-06-05arXiv.orgCitations: 1