Reward Weighted Classifier-Free Guidance as Policy Improvement in Autoregressive Models

📅 2026-04-16
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Autoregressive models typically require retraining to align with new reward functions when they change, leading to inefficiency. This work proposes Reward-Conditioned Classifier-Free Guidance (RCFG), which formalizes reward-guided generation as a policy improvement operator for the first time, enabling flexible optimization of arbitrary reward functions at test time without retraining. By integrating autoregressive modeling, policy distillation, and reinforcement learning, RCFG supports zero-shot reward adaptation and serves as an effective warm-start to accelerate subsequent reinforcement learning convergence. Evaluated on molecular design tasks, RCFG demonstrates substantial improvements in both test-time reward optimization capability and training efficiency.

Technology Category

Machine Learning: Deep Generative Models & AutoencodersComputer Vision: Diffusion Models for VisionNatural Language Processing: Generation

Application Category

Search and Retrieval-Augmented AI: Retrieval-Augmented Generation (RAG) and multi-modal RAGUser Modeling, Personalization and Recommendation: Fairness-aware retrieval and rankingResponsible Web: Machine-in-the-loop, human agency and autonomy
📝 Abstract
Consider an auto-regressive model that produces outputs x (e.g., answers to questions, molecules) each of which can be summarized by an attribute vector y (e.g., helpfulness vs. harmlessness, or bio-availability vs. lipophilicity). An arbitrary reward function r(y) encodes tradeoffs between these properties. Typically, tilting the model's sampling distribution to increase this reward is done at training time via reinforcement learning. However, if the reward function changes, re-alignment requires re-training. In this paper, we show that a reward weighted classifier-free guidance (RCFG) can act as a policy improvement operator in this setting, approximating tilting the sampling distribution by the Q function. We apply RCFG to molecular generation, demonstrating that it can optimize novel reward functions at test time. Finally, we show that using RCFG as a teacher and distilling into the base policy to serve as a warm start significantly speeds up convergence for standard RL.
Problem

Research questions and friction points this paper is trying to address.

reward function
autoregressive models
policy improvement
test-time optimization
model alignment
Innovation

Methods, ideas, or system contributions that make the work stand out.

Reward Weighted Classifier-Free Guidance
Policy Improvement
Autoregressive Models
Test-Time Optimization
Policy Distillation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
A
Alexander Peysakhovich
Sutter Hill Ventures
W
William Berman
Sutter Hill Ventures