CT-OPD: Counterfactual Trace On-Policy Distillation for Diffusion Vision-Language Models
This study addresses the misalignment between preset masks and model decisions during the post-training of diffusion vision-language models, as well as the sparse feedback inherent in reinforcement learning. To overcome these challenges, we propose a counterfactual trajectory online distillation method. Specifically, this approach extracts masks from the student's current trajectory to reconstruct the teacher's endpoint states, leveraging the teacher's complete response to provide coherent token-level supervision that precisely aligns training signals with the model's revealed decisions. Experimental results demonstrate that our method yields an average improvement of 9.80 points across nine benchmarks. It significantly enhances multimodal understanding and reasoning capabilities while simultaneously improving image generation quality within a unified architecture.