🤖 AI Summary
This study addresses the problem of inaccurate reward estimation caused by LLM reviewer bias in active preference learning, where conventional calibration fails to eliminate residual bias from distribution shift. We propose the NAOD strategy, which explicitly incorporates reviewer bias into acquisition design. By establishing an asymptotic minimax lower bound on policy risk and constructing an optimal estimator, NAOD integrates joint estimation with nuisance adjustment and optimizes comparison selection via the Frank-Wolfe algorithm to prioritize target-relevant information. Theoretically, we reveal a reversal effect wherein representation error undermines the advantages of oracle-based designs. Empirically, experiments on Chatbot Arena data demonstrate that NAOD reduces the average regret of surrogate policies by 29.1%, significantly outperforming existing methods while improving human preference prediction accuracy.
📝 Abstract
Learning from human preferences is central to large language model (LLM) alignment, but human preference annotation is costly. Active preference learning reduces this cost by selecting informative comparisons, and LLM judges can provide additional scalable feedback. However, the preferences of the judges may deviate from those of the target human population. Even after calibration on trusted reference data, active acquisition can shift the comparison distribution and expose residual judge bias. We therefore incorporate judge deviations into the acquisition design rather than relying on a separate calibration stage. Under joint estimation, comparisons that appear highly informative about the reward may also reflect judge bias and therefore provide less information about human preferences. To address this issue, we propose Nuisance-Adjusted Optimal Design (NAOD), a comparison-selection strategy that prioritizes policy-relevant target information after nuisance adjustment and uses the Frank-Wolfe algorithm for optimization. Theoretically, we establish a sharp conditional local asymptotic minimax lower bound on policy risk and construct an estimator that attains it. We further characterize the finite-sample cost of learning the nuisance representation and show that representation error can reverse an oracle design advantage. Finally, we validate these predictions experimentally and evaluate NAOD on Chatbot Arena data across 17 judges, 15 budget configurations, and 15 random cluster-level splits. NAOD reduces the mean regret of proxy policy by 29.1% relative to a matched target-information design, outperforms existing methods, and improves human-preference prediction on held-out data.