🤖 AI Summary
Existing alignment methods—such as RLHF, human/internal feedback integration, test-time scaling, and diffusion guidance—lack a unified theoretical foundation, leading to instability (e.g., reward hacking, policy optimization divergence) and inflexibility in behavior control.
Method: We establish the intrinsic equivalence among these paradigms, showing that soft-optimal N-sampling, resampling-based guidance, and reward modeling all instantiate implicit policy optimization. Building on this insight, we propose a novel resampling-based alignment framework that operates entirely at inference time: it dynamically reweights or resamples diffusion trajectories using heterogeneous feedback signals—explicit human ratings and implicit model self-assessments—without explicit RL training.
Contribution/Results: Our approach eliminates reliance on unstable policy network updates and reward modeling pitfalls while preserving generation quality. It achieves superior controllability, robustness, and adaptability across diverse alignment objectives. Crucially, this work provides the first systematic theoretical unification of mainstream post-training alignment techniques under a coherent implicit optimization lens.
📝 Abstract
In this note, we reflect on several fundamental connections among widely used post-training techniques. We clarify some intimate connections and equivalences between reinforcement learning with human feedback, reinforcement learning with internal feedback, and test-time scaling (particularly soft best-of-$N$ sampling), while also illuminating intrinsic links between diffusion guidance and test-time scaling. Additionally, we introduce a resampling approach for alignment and reward-directed diffusion models, sidestepping the need for explicit reinforcement learning techniques.