Unified Trajectory Matching Policy Optimization: Diverse T2I Generation and VLA Generalization
This study addresses the lack of diversity and limited generalization caused by policy mode collapse during reinforcement learning post-training. To this end, we propose Uni-TMPO, a unified framework that replaces conventional reward maximization with trajectory matching optimization. By introducing forward KL divergence for distribution alignment, combined with a coarse-to-fine scheduler and feedback-conditioned sampling, our method enables unified training across text-to-image (T2I) and vision-language-action (VLA) models. Experimental results demonstrate that Uni-TMPO significantly improves T2I reward scores and VLA task success rates while effectively balancing generation quality, diversity, and computational efficiency. Furthermore, the framework enhances cross-scenario generalization capabilities, and its efficacy in generating diverse policies is validated through real-world robotic deployment.