🤖 AI Summary
This study addresses the lack of diversity and limited generalization caused by policy mode collapse during reinforcement learning post-training. To this end, we propose Uni-TMPO, a unified framework that replaces conventional reward maximization with trajectory matching optimization. By introducing forward KL divergence for distribution alignment, combined with a coarse-to-fine scheduler and feedback-conditioned sampling, our method enables unified training across text-to-image (T2I) and vision-language-action (VLA) models. Experimental results demonstrate that Uni-TMPO significantly improves T2I reward scores and VLA task success rates while effectively balancing generation quality, diversity, and computational efficiency. Furthermore, the framework enhances cross-scenario generalization capabilities, and its efficacy in generating diverse policies is validated through real-world robotic deployment.
📝 Abstract
Reward-maximizing reinforcement learning (RL) is widely used to post-train stochastic diffusion and flow policies for text-to-image (T2I) generation. However, reward-maximizing RL causes policy mode collapse even under reference KL or entropy regularization, reducing the policy to a single high-reward mode. In T2I, this produces similar images and reward hacking. When extended to vision-language-action (VLA) models, the same collapse removes alternative successful strategies and weakens task and scene generalization. To address this limitation, we introduce Unified Trajectory Matching Policy Optimization (Uni-TMPO), a unified RL post-training framework for diffusion and flow policies. First, Uni-TMPO converts standardized rewards into a target distribution within each trajectory group and derives the policy distribution from trajectory log probabilities. Then, forward Kullback-Leibler optimization matches the two distributions instead of maximizing expected reward. A progress-conditioned coarse-to-fine scheduler efficiently constructs T2I trajectories. Within the unified framework, feedback-conditioned sampling uses updated observations to construct VLA trajectories. Extensive experiments show that Uni-TMPO achieves higher T2I rewards and VLA ID success rates than the strongest baselines. More importantly, it achieves the best T2I reward-diversity-efficiency trade-off and VLA generalization to held-out tasks and scenes, while real-robot evaluation demonstrates the value of multiple action strategies when the higher-reward target is blocked.