Pointwise or Pairwise: When Do Pairwise Losses Help Reward Learning, Provably?

๐Ÿ“… 2026-09-29
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This study delineates the conditions under which pairwise losses outperform pointwise losses in reward learning, as well as their failure regimes. Under a grouped offline contextual bandit setting, we develop a unified local analysis framework to compare value regression with value difference regression (VDR), providing the first finite-sample theoretical guarantees and revealing a biasโ€“variance tradeoff governed by feature geometry. Through semiparametric modeling, function class analysis, and regret bound derivations, we demonstrate that VDR eliminates misspecification terms and optimizes estimation error, yet may inflate variance when model misspecification is minimal. All theoretical findings are corroborated empirically. This work systematically clarifies the theoretical applicability conditions for pairwise losses in reward modeling.
๐Ÿ“ Abstract
Pairwise losses are increasingly used for reward learning even when pointwise rewards are observed, with mixed empirical results. When and why do pairwise losses outperform pointwise losses? We study this question in a grouped offline contextual-bandit setting allowing multiple actions per context, capturing many reward learning scenarios. We compare Value Regression (VR), which regresses observed rewards pointwise, with Value Difference Regression (VDR), which regresses reward differences between a pair of actions sampled under the same context. We consider a semiparametric model where the mean reward is the sum of a learnable action-dependent component and an arbitrary context-dependent yet action-independent nuisance, capturing context-specific disturbances. Using a unified localized analysis, we prove finite-sample regression guarantees for finite and linear function classes and translate them into offline-regret bounds. For finite classes, VDR eliminates the misspecification term in the VR bound and improves a reward-scale-dependent error term by averaging over actions within each context, a benefit absent from the corresponding VR term. For linear classes, neither method uniformly dominates: within-context differencing removes nuisance-induced bias but may increase estimation variance relative to using absolute rewards when the misspecification is sufficiently low. This yields a feature geometry-dependent bias-variance tradeoff, which we corroborate with numerical experiments.
Problem

Research questions and friction points this paper is trying to address.

Reward Learning
Pairwise Losses
Pointwise Losses
Contextual Bandits
Offline Learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Pairwise Loss
Reward Learning
Contextual Bandits
Semiparametric Model
Bias-Variance Tradeoff
๐Ÿ”Ž Similar Papers
No similar papers found.