Score
Designs, builds, and evaluates models, architectures, and end‑to‑end systems that generate novel content across two or more data modalities (such as text, images, audio, and video), including conditional and joint generation and mechanisms for cross‑modal alignment and consistency. Works include creating generation algorithms, training and inference pipelines, datasets and evaluation metrics focused on multimodal quality, coherence, and controllability.
This survey addresses the growing complexity and fragmentation in AI-generated content (AIGC) research across diverse modalities. We systematically review generative methods and cross-modal translation techniques—including text-to-image, audio-to-video, and others—spanning seven modalities: text, image, video, 3D shape/scene/portrait/motion, and audio. Methodologically, we establish the first unified analytical framework covering all modalities, grounded in foundational architectures: GANs, VAEs, diffusion models, autoregressive Transformers, and multimodal alignment paradigms (e.g., CLIP, Flux). Our key contributions include: (i) a novel taxonomy of cross-modal generation paradigms; (ii) a horizontal multimodal comparison framework; (iii) synthesis of 120+ representative works; and (iv) a consolidated analysis of datasets, evaluation metrics, shared challenges, performance bottlenecks, and a comparative performance table. The survey provides systematic guidance for AIGC technology selection, benchmark development, and future research directions.
Current multimodal generative models face two key bottlenecks in design assistance: insufficient comprehension of ambiguous instructions and difficulty maintaining both content consistency and creativity under reference guidance. To address these, we propose WeGen—the first unified architecture enabling bidirectional generation-understanding co-evolution. It integrates dynamical alignment via interleaved sequence modeling, consistency-aware generation, and prompt self-rewriting to support interactive, iterative multimodal creation. Built upon multimodal sequence modeling, WeGen leverages foundation-model–self-annotated dynamical datasets and interleaved object-dynamics representations, enabling controllable refinement while preserving user-satisfying content. Experiments demonstrate that WeGen achieves state-of-the-art performance on visual generation benchmarks, significantly improving creativity, reference fidelity, and user controllability—validating its effectiveness as an efficient, intuitive design collaborator.
Existing AIGC models are often confined to single modalities or specific scenarios, struggling to generate complex multimodal content that seamlessly integrates audio, video, and text in an end-to-end manner. This work proposes a multimodal agent system grounded in skill acquisition theory to guide data construction and training. The approach features a two-stage planning optimization strategy—comprising autocorrelation modeling and preference alignment—and a three-stage fine-tuning pipeline involving base training, successful-plan fine-tuning, and preference optimization. Experimental results demonstrate that the proposed method significantly outperforms current state-of-the-art models, achieving notable improvements in both generation quality and alignment with human preferences.
To address challenges in multimodal narrative storytelling—including poor structural control, weak cross-modal consistency, and coarse-grained editing—this paper proposes a graph-node-based multimodal content generation system. Methodologically, it introduces a node-centric editing framework that maps text, images, audio, and video to editable graph nodes; employs a task-selection agent for dynamic generation orchestration; and integrates multimodal large language models, context-aware generation, and natural language understanding to enable node-level precise editing, automatic parallel narrative branching, and cross-modal co-evolution. Contributions include: (1) the first fine-grained, interpretable, human-in-the-loop multimodal narrative iterative generation system; (2) significant performance gains on story outline generation; and (3) user studies demonstrating a 42% improvement in editing efficiency, enhanced creative flexibility, and validated alignment between controllability and effectiveness.
Existing unified multimodal models are largely restricted to unidirectional, single-modality generation, failing to support sequence-level interleaved co-generation of text and images. To address this, we propose Mogao—the first unified foundation model capable of generating arbitrarily long, interleaved text-image sequences. Its core innovations include: (i) interleaved rotary position encoding, (ii) a dual-visual-encoder architecture, (iii) a multimodal classifier-free guidance mechanism, and (iv) a joint training paradigm integrating causal modeling with diffusion priors. Mogao achieves zero-shot image editing and compositional generation—emergent capabilities previously unattainable in unified models. It establishes new state-of-the-art performance on multimodal understanding and text-to-image synthesis. Moreover, it generates high-fidelity, semantically coherent interleaved sequences and significantly improves the quality of complex edits and compositional generation.
Current text-to-video models suffer from significant deficiencies in physical plausibility, photorealistic lighting, camera motion, and temporal coherence, limiting their applicability to cinematic-grade synthesis. To address this, we propose the first multi-agent VLM framework tailored for high-fidelity 3D video generation, featuring decoupled Director, Programmer, and Reviewer agents. Our method decomposes the synthesis task, automatically generates Blender scripting code, and performs iterative optimization guided by vision-language feedback—enabling end-to-end, interpretable, and editable video generation. Deeply integrating cinematographic knowledge with a closed-loop 3D rendering pipeline, it produces high-fidelity videos fully aligned with textual prompts—without manual intervention. Experiments demonstrate superior performance over leading commercial models across five video quality and instruction-following metrics. User studies further confirm substantial improvements: +28.6% in physical plausibility, +31.2% in temporal consistency, and higher overall quality scores.
This work addresses the limitations of existing approaches that predominantly rely on non-native late-fusion architectures and lack a systematic definition or unified framework for native multimodal modeling. The paper proposes a formal roadmap for Native Multimodal Modeling (NMM), introducing, for the first time, a rigorous formulation of “architectural nativeness.” Leveraging an input–output duality perspective, it categorizes models into three types: Multi-to-Text, Multi-to-Target, and Multi-to-Multi. A full-stack, industrial-grade NMM framework is developed, encompassing data governance, early/mid-stage fusion, end-to-end training, and deployment, all realized through a unified Transformer architecture that enables symbiotic cross-modal understanding and generation. This paradigm demonstrates superior performance and strong scalability across multiple tasks, offering a clear pathway toward truly native multimodal models.
Existing interleaved multimodal generative models struggle to achieve high-quality text-image interleaved generation due to scarce training data and limited capabilities of underlying foundation models. This work proposes DuoGen, a framework that leverages a large-scale, high-quality instruction-tuning dataset and integrates a multimodal large language model (MLLM) with a video-pretrained diffusion Transformer (DiT). DuoGen employs a two-stage decoupled strategy to jointly optimize comprehension and generation capabilities, eliminating the need for costly unimodal pretraining and enabling flexible selection of foundation models. Furthermore, it establishes the first comprehensive evaluation benchmark tailored for interleaved generation. Experiments demonstrate that DuoGen significantly outperforms existing open-source models in text quality, image fidelity, and text-image alignment, achieving state-of-the-art performance in both text-to-image generation and image editing within a unified architecture.
This work addresses the limitations of existing video generation models, which often neglect audio and rely on cascaded pipelines, leading to high computational costs, error propagation, and audio-visual desynchronization. To overcome these challenges, we propose MOVA—the first open-source, end-to-end image-and-text-to-video-and-audio (IT2VA) generation model. Built upon a 32B-parameter Mixture-of-Experts architecture (with 18B activated per forward pass), MOVA supports LoRA-based fine-tuning, efficient inference, and prompt enhancement. It simultaneously generates semantically aligned, high-fidelity video and audio, including lip-synced speech, contextually appropriate sound effects, and background music. By releasing the model weights and a complete toolchain, this work aims to advance research in joint audio-visual synthesis.