Self-correction Optimization for Interleaved Multimodal Generation

πŸ“… 2026-10-07
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the challenges of visual subject inconsistency, poor temporal coherence, and insufficient physical plausibility in multimodal interleaved image-text generation. To this end, we propose a training-free self-correcting optimization framework. Building upon classifier-free guidance (CFG), our method introduces a complementary constraint mechanism that jointly governs novel event incorporation and state preservation, achieving self-correction by imposing minimal modifications to the guided updates without requiring additional training or computational overhead. Experimental results demonstrate that the proposed approach significantly enhances both the temporal coherence and physical realism of generated content across complex benchmarks. Furthermore, it successfully generalizes to long-horizon video generation tasks, such as robotic manipulation, highlighting its broad applicability and effectiveness.
πŸ“ Abstract
Multimodal large language models (MLLMs) have made significant progress in visual understanding and generation. However, generating interleaved image--text content remains challenging, as it requires tightly integrated multimodal understanding and generation capabilities. Although existing MLLMs provide promising solutions, most rely on additional training with augmented data, which is computationally expensive and remains limited in preserving visual subjects, temporal consistency, and physical plausibility. In this work, we propose self-correction optimization (SCO), an effective training-free method for consistent interleaved generation. SCO treats the classifier-free guidance update as a reference and performs minimal self-correction under two complementary constraints, including new-event and state-preserving constraints. Specifically, the new-event constraint promotes temporal consistency across image--text sequences, while the state-preserving constraint maintains the coherence of visual subjects throughout subsequent generation steps. Experiments on challenging interleaved multimodal generation benchmarks demonstrate significant improvements in temporal coherence and visual-subject preservation. Furthermore, SCO can be extended to video generation and improves the modeling of physically grounded processes, including robot manipulation and long-horizon handcrafting.
Problem

Research questions and friction points this paper is trying to address.

interleaved multimodal generation
multimodal large language models
temporal consistency
visual-subject preservation
physical plausibility
Innovation

Methods, ideas, or system contributions that make the work stand out.

Self-correction Optimization
Interleaved Multimodal Generation
Training-free
Classifier-free Guidance
Temporal Consistency
πŸ”Ž Similar Papers