๐ค AI Summary
This study addresses whether KL-constrained reinforcement learning and Best-of-n sampling are theoretically equivalent for language model alignment when output sequences exhibit memory. Departing from the conventional independent and identically distributed assumption, this work integrates probability theory, information theory, and Markov process modeling to extend the asymptotic consistency analysis of these two mainstream alignment methods to Markov chain settings. The primary contribution is establishing, for the first time, a theory of asymptotic equivalence between these alignment approaches under memory-dependent language models. Specifically, it proves the asymptotic closeness of the distributions induced by both algorithms under Markovian outputs. Furthermore, for sequence length m=1, it derives the necessary and sufficient conditions for the KullbackโLeibler divergence to vanish, thereby providing a complete characterization of zero-divergence scenarios.
๐ Abstract
Language model (LM) alignment broadly aims to perturb a given LM $Q$ into an aligned LM $q$ such that i) the outputs produced by $q$ and $Q$ are 'close' in probability, ii) $q$ has a higher expected reward than $Q$. Two common techniques for LM alignment are: KL-constrained RL, which requires knowledge of the LM distribution and is computationally expensive, and the best-of-$n$ algorithm, which requires only sampling from the LM. The work of Yang et al. established asymptotic closeness between the distributions produced by the two alignment methods for an $m$--length i.i.d. token sequence output by the LM, in the limit as $m$ increases to infinity. However, the i.i.d. assumption is not representative of practical LMs, whose output sequences often have memory. In this paper, we extend the asymptotic closeness result to the case when the $m$--length token sequence outputted by the LM is Markovian. Further, for finite-length output sequences -- particularly, when $m=1$ -- we provide a complete characterization of LM distributions and reward functions for which the KL-divergence between the distributions produced by the two alignment methods is zero -- a question first posed in Yang et al.