π€ AI Summary
This study addresses the challenges of visual subject inconsistency, poor temporal coherence, and insufficient physical plausibility in multimodal interleaved image-text generation. To this end, we propose a training-free self-correcting optimization framework. Building upon classifier-free guidance (CFG), our method introduces a complementary constraint mechanism that jointly governs novel event incorporation and state preservation, achieving self-correction by imposing minimal modifications to the guided updates without requiring additional training or computational overhead. Experimental results demonstrate that the proposed approach significantly enhances both the temporal coherence and physical realism of generated content across complex benchmarks. Furthermore, it successfully generalizes to long-horizon video generation tasks, such as robotic manipulation, highlighting its broad applicability and effectiveness.
π Abstract
Multimodal large language models (MLLMs) have made significant progress in visual understanding and generation. However, generating interleaved image--text content remains challenging, as it requires tightly integrated multimodal understanding and generation capabilities. Although existing MLLMs provide promising solutions, most rely on additional training with augmented data, which is computationally expensive and remains limited in preserving visual subjects, temporal consistency, and physical plausibility. In this work, we propose self-correction optimization (SCO), an effective training-free method for consistent interleaved generation. SCO treats the classifier-free guidance update as a reference and performs minimal self-correction under two complementary constraints, including new-event and state-preserving constraints. Specifically, the new-event constraint promotes temporal consistency across image--text sequences, while the state-preserving constraint maintains the coherence of visual subjects throughout subsequent generation steps. Experiments on challenging interleaved multimodal generation benchmarks demonstrate significant improvements in temporal coherence and visual-subject preservation. Furthermore, SCO can be extended to video generation and improves the modeling of physically grounded processes, including robot manipulation and long-horizon handcrafting.