🤖 AI Summary
This study addresses the problem of sequence probability decay with length in autoregressive generation for large language models by formulating text revision as a game-theoretic process. Specifically, token positions are modeled as players, the vocabulary serves as the action space, and utility functions are defined via log-conditional probabilities. This work introduces Nash decoding, a novel algorithm that optimizes joint generation through rapid convergence to an ε-Nash equilibrium. Theoretically, we prove that this equilibrium yields exponential likelihood improvements for long sequences. Empirically, without requiring fine-tuning, the proposed method achieves F1 and ROUGE scores on multiple question-answering benchmarks that significantly surpass those of autoregressive models with eighteen times more parameters.
📝 Abstract
Text revision has become an integral component of large language models. This paper formulates revision such that it admits a Nash equilibrium: Token positions are players, vocabulary items are actions, and each player's utility is the language model's log conditional probability. We motivate the revision by showing that Nash equilibria can have exponentially higher likelihood than autoregressive outputs as the sequence length grows. We further propose Nash decoding, an algorithm that reaches an $\varepsilon$-Nash equilibrium in $O(1/\varepsilon)$ time given access to the joint probability of tokens conditioned on a prompt. In practice, we run Nash decoding using conditional probability estimates from large language models and evaluate the resulting equilibria on question-answering benchmarks. On CLAPNQ, PubMedQA, and CoQA, Nash equilibria obtained from masked language models achieve higher F1 and ROUGE scores than autoregressive models up to $18\times$ larger, without any fine-tuning or retraining, at the cost of additional test-time computation.