Connections between reinforcement learning with feedback,test-time scaling, and diffusion guidance: An anthology

📅 2025-09-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Existing alignment methods—such as RLHF, human/internal feedback integration, test-time scaling, and diffusion guidance—lack a unified theoretical foundation, leading to instability (e.g., reward hacking, policy optimization divergence) and inflexibility in behavior control. Method: We establish the intrinsic equivalence among these paradigms, showing that soft-optimal N-sampling, resampling-based guidance, and reward modeling all instantiate implicit policy optimization. Building on this insight, we propose a novel resampling-based alignment framework that operates entirely at inference time: it dynamically reweights or resamples diffusion trajectories using heterogeneous feedback signals—explicit human ratings and implicit model self-assessments—without explicit RL training. Contribution/Results: Our approach eliminates reliance on unstable policy network updates and reward modeling pitfalls while preserving generation quality. It achieves superior controllability, robustness, and adaptability across diverse alignment objectives. Crucially, this work provides the first systematic theoretical unification of mainstream post-training alignment techniques under a coherent implicit optimization lens.

Technology Category

Search and Optimization: Sampling/Simulation-based SearchComputer Vision: Diffusion Models for VisionMachine Learning: Imitation Learning & Inverse Reinforcement Learning

Application Category

Search and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingUser Modeling, Personalization and Recommendation: Fairness-aware retrieval and rankingGraph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphs
📝 Abstract
In this note, we reflect on several fundamental connections among widely used post-training techniques. We clarify some intimate connections and equivalences between reinforcement learning with human feedback, reinforcement learning with internal feedback, and test-time scaling (particularly soft best-of-$N$ sampling), while also illuminating intrinsic links between diffusion guidance and test-time scaling. Additionally, we introduce a resampling approach for alignment and reward-directed diffusion models, sidestepping the need for explicit reinforcement learning techniques.
Problem

Research questions and friction points this paper is trying to address.

Connecting reinforcement learning feedback and test-time scaling
Exploring equivalences between human and internal feedback methods
Introducing resampling for diffusion models without reinforcement learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Connects reinforcement learning with feedback
Links diffusion guidance to test-time scaling
Introduces resampling for alignment without RL
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Y
Yuchen Jiao
Department of Statistics and Data Science, Chinese University of Hong Kong.
Y
Yuxin Chen
Department of Statistics and Data Science, the Wharton School, University of Pennsylvania.
G
Gen Li
Department of Statistics and Data Science, Chinese University of Hong Kong.