Unified Trajectory Matching Policy Optimization: Diverse T2I Generation and VLA Generalization

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of diversity and limited generalization caused by policy mode collapse during reinforcement learning post-training. To this end, we propose Uni-TMPO, a unified framework that replaces conventional reward maximization with trajectory matching optimization. By introducing forward KL divergence for distribution alignment, combined with a coarse-to-fine scheduler and feedback-conditioned sampling, our method enables unified training across text-to-image (T2I) and vision-language-action (VLA) models. Experimental results demonstrate that Uni-TMPO significantly improves T2I reward scores and VLA task success rates while effectively balancing generation quality, diversity, and computational efficiency. Furthermore, the framework enhances cross-scenario generalization capabilities, and its efficacy in generating diverse policies is validated through real-world robotic deployment.
📝 Abstract
Reward-maximizing reinforcement learning (RL) is widely used to post-train stochastic diffusion and flow policies for text-to-image (T2I) generation. However, reward-maximizing RL causes policy mode collapse even under reference KL or entropy regularization, reducing the policy to a single high-reward mode. In T2I, this produces similar images and reward hacking. When extended to vision-language-action (VLA) models, the same collapse removes alternative successful strategies and weakens task and scene generalization. To address this limitation, we introduce Unified Trajectory Matching Policy Optimization (Uni-TMPO), a unified RL post-training framework for diffusion and flow policies. First, Uni-TMPO converts standardized rewards into a target distribution within each trajectory group and derives the policy distribution from trajectory log probabilities. Then, forward Kullback-Leibler optimization matches the two distributions instead of maximizing expected reward. A progress-conditioned coarse-to-fine scheduler efficiently constructs T2I trajectories. Within the unified framework, feedback-conditioned sampling uses updated observations to construct VLA trajectories. Extensive experiments show that Uni-TMPO achieves higher T2I rewards and VLA ID success rates than the strongest baselines. More importantly, it achieves the best T2I reward-diversity-efficiency trade-off and VLA generalization to held-out tasks and scenes, while real-robot evaluation demonstrates the value of multiple action strategies when the higher-reward target is blocked.
Problem

Research questions and friction points this paper is trying to address.

policy mode collapse
reinforcement learning
text-to-image generation
vision-language-action models
generalization
Innovation

Methods, ideas, or system contributions that make the work stand out.

Trajectory Matching Policy Optimization
Forward KL Divergence
Mode Collapse Mitigation
Diffusion and Flow Policies
Vision-Language-Action Models
🔎 Similar Papers
2024-09-12IEEE Transactions on Automation Science and EngineeringCitations: 2