Mutually Adversarial Self-Training with Evolving Data for Unified Multimodal Models

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation that the generation and understanding branches of unified multimodal models lack adversarial challenges during collaborative self-training, thereby constraining performance gains. To overcome this, we propose MATE, a novel reinforcement learning-based mutually adversarial self-training framework. In MATE, the two branches alternately serve as challenger and solver, engaging in self-play within the data space. The model leverages its own outputs to construct dynamically evolving adversarial data, incorporating consistency-based filtering to ensure quality without requiring an independently trained adversary model. Experiments on Janus-Pro-1B demonstrate that MATE significantly improves performance on GenEval, DPG-Bench, and multiple understanding benchmarks, while effectively enhancing image-text cycle consistency.
📝 Abstract
Unified multimodal models (UMMs) combine image generation and visual understanding in a shared backbone. Since generation and understanding are inverse tasks, recent studies self-train UMMs by letting the two branches cooperatively supervise each other. We introduce MATE (Mutually Adversarial self-Training with Evolving data), a reinforcement-learning-based post-training framework in which the two branches instead challenge each other, and the challenges evolve as the model trains. MATE lets generation and understanding take turns to be challenger and solver. Given an image, the understanding branch proposes several candidate descriptions that the generation branch must turn back into similar images, and vice versa. The candidates are screened for consistency with the image or prompt they were proposed from, and the solver is trained on the candidate it handles worst. The adversary thus comes from the model's own outputs, and no separate adversary is trained. Moreover, the candidates that defeat one branch become the sources of the next challenges to the other in the next epoch, which keeps the challenges evolving with the model and turns the training into self-play in data space. On Janus-Pro-1B, MATE improves GenEval by 2.4 points, DPG-Bench by 1.7 points, and the average over nine understanding benchmarks by 0.7 points, while strengthening consistency across repeated image-text cycles.
Problem

Research questions and friction points this paper is trying to address.

Unified Multimodal Models
Self-Training
Image Generation
Visual Understanding
Adversarial Training
Innovation

Methods, ideas, or system contributions that make the work stand out.

Mutually Adversarial Self-Training
Unified Multimodal Models
Reinforcement Learning
Self-play in Data Space
Evolving Data
🔎 Similar Papers
No similar papers found.