multimodal generation

Designs, builds, and evaluates models, architectures, and end‑to‑end systems that generate novel content across two or more data modalities (such as text, images, audio, and video), including conditional and joint generation and mechanisms for cross‑modal alignment and consistency. Works include creating generation algorithms, training and inference pipelines, datasets and evaluation metrics focused on multimodal quality, coherence, and controllability.

multimodalgeneration

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-1.14
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$238K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

WeGen: A Unified Model for Interactive Multimodal Generation as We Chat

Mar 03, 2025
ZH
Zhipeng Huang
🏛️ University of Science and Technology of China | Shanghai Jiao Tong University | Tencent Inc | Galbot | Chinese Academy of Sciences

Current multimodal generative models face two key bottlenecks in design assistance: insufficient comprehension of ambiguous instructions and difficulty maintaining both content consistency and creativity under reference guidance. To address these, we propose WeGen—the first unified architecture enabling bidirectional generation-understanding co-evolution. It integrates dynamical alignment via interleaved sequence modeling, consistency-aware generation, and prompt self-rewriting to support interactive, iterative multimodal creation. Built upon multimodal sequence modeling, WeGen leverages foundation-model–self-annotated dynamical datasets and interleaved object-dynamics representations, enabling controllable refinement while preserving user-satisfying content. Experiments demonstrate that WeGen achieves state-of-the-art performance on visual generation benchmarks, significantly improving creativity, reference fidelity, and user controllability—validating its effectiveness as an efficient, intuitive design collaborator.

Enhances multimodal generation for less detailed instructionsImproves creativity and diversity in visual content generationMaintains consistency with user references during iterative generation

Existing AIGC models are often confined to single modalities or specific scenarios, struggling to generate complex multimodal content that seamlessly integrates audio, video, and text in an end-to-end manner. This work proposes a multimodal agent system grounded in skill acquisition theory to guide data construction and training. The approach features a two-stage planning optimization strategy—comprising autocorrelation modeling and preference alignment—and a three-stage fine-tuning pipeline involving base training, successful-plan fine-tuning, and preference optimization. Experimental results demonstrate that the proposed method significantly outperforms current state-of-the-art models, achieving notable improvements in both generation quality and alignment with human preferences.

agent-based systemsAIGCcontent creation

Node-Based Editing for Multimodal Generation of Text, Audio, Image, and Video

Nov 05, 2025
AH
Alexander Htet Kyaw
🏛️ Massachusetts Institute of Technology | Microsoft Research

To address challenges in multimodal narrative storytelling—including poor structural control, weak cross-modal consistency, and coarse-grained editing—this paper proposes a graph-node-based multimodal content generation system. Methodologically, it introduces a node-centric editing framework that maps text, images, audio, and video to editable graph nodes; employs a task-selection agent for dynamic generation orchestration; and integrates multimodal large language models, context-aware generation, and natural language understanding to enable node-level precise editing, automatic parallel narrative branching, and cross-modal co-evolution. Contributions include: (1) the first fine-grained, interpretable, human-in-the-loop multimodal narrative iterative generation system; (2) significant performance gains on story outline generation; and (3) user studies demonstrating a 42% improvement in editing efficiency, enhanced creative flexibility, and validated alignment between controllability and effectiveness.

Creating multimodal narratives integrating text, audio, images and videoEnabling iterative refinement of story structure through node editingMaintaining narrative consistency across multiple nodes and modalities

Mogao: An Omni Foundation Model for Interleaved Multi-Modal Generation

May 08, 2025
CL
Chao Liao
🏛️ ByteDance | Seed

Existing unified multimodal models are largely restricted to unidirectional, single-modality generation, failing to support sequence-level interleaved co-generation of text and images. To address this, we propose Mogao—the first unified foundation model capable of generating arbitrarily long, interleaved text-image sequences. Its core innovations include: (i) interleaved rotary position encoding, (ii) a dual-visual-encoder architecture, (iii) a multimodal classifier-free guidance mechanism, and (iv) a joint training paradigm integrating causal modeling with diffusion priors. Mogao achieves zero-shot image editing and compositional generation—emergent capabilities previously unattainable in unified models. It establishes new state-of-the-art performance on multimodal understanding and text-to-image synthesis. Moreover, it generates high-fidelity, semantically coherent interleaved sequences and significantly improves the quality of complex edits and compositional generation.

Advancing unified models with large-scale joint training strategyEnabling interleaved multi-modal generation via causal approachIntegrating autoregressive and diffusion models for text-image synthesis

Kubrick: Multimodal Agent Collaborations for Synthetic Video Generation

Aug 19, 2024
LH
Liu He
🏛️ Purdue University | Baidu

Current text-to-video models suffer from significant deficiencies in physical plausibility, photorealistic lighting, camera motion, and temporal coherence, limiting their applicability to cinematic-grade synthesis. To address this, we propose the first multi-agent VLM framework tailored for high-fidelity 3D video generation, featuring decoupled Director, Programmer, and Reviewer agents. Our method decomposes the synthesis task, automatically generates Blender scripting code, and performs iterative optimization guided by vision-language feedback—enabling end-to-end, interpretable, and editable video generation. Deeply integrating cinematographic knowledge with a closed-loop 3D rendering pipeline, it produces high-fidelity videos fully aligned with textual prompts—without manual intervention. Experiments demonstrate superior performance over leading commercial models across five video quality and instruction-following metrics. User studies further confirm substantial improvements: +28.6% in physical plausibility, +31.2% in temporal consistency, and higher overall quality scores.

Address improper motion and consistency in text-to-video generationAutomate synthetic video creation via VLM agent collaborationReduce manual CGI editing in film industry workflows

Latest Papers

What's happening recently
View more

This work addresses the limitations of existing approaches that predominantly rely on non-native late-fusion architectures and lack a systematic definition or unified framework for native multimodal modeling. The paper proposes a formal roadmap for Native Multimodal Modeling (NMM), introducing, for the first time, a rigorous formulation of “architectural nativeness.” Leveraging an input–output duality perspective, it categorizes models into three types: Multi-to-Text, Multi-to-Target, and Multi-to-Multi. A full-stack, industrial-grade NMM framework is developed, encompassing data governance, early/mid-stage fusion, end-to-end training, and deployment, all realized through a unified Transformer architecture that enables symbiotic cross-modal understanding and generation. This paradigm demonstrates superior performance and strong scalability across multiple tasks, offering a clear pathway toward truly native multimodal models.

architectural nativitymultimodal architecturemultimodal fusion

Existing interleaved multimodal generative models struggle to achieve high-quality text-image interleaved generation due to scarce training data and limited capabilities of underlying foundation models. This work proposes DuoGen, a framework that leverages a large-scale, high-quality instruction-tuning dataset and integrates a multimodal large language model (MLLM) with a video-pretrained diffusion Transformer (DiT). DuoGen employs a two-stage decoupled strategy to jointly optimize comprehension and generation capabilities, eliminating the need for costly unimodal pretraining and enabling flexible selection of foundation models. Furthermore, it establishes the first comprehensive evaluation benchmark tailored for interleaved generation. Experiments demonstrate that DuoGen significantly outperforms existing open-source models in text quality, image fidelity, and text-image alignment, achieving state-of-the-art performance in both text-to-image generation and image editing within a unified architecture.

base model capacitygeneral-purpose generationimage-text alignment

This work addresses the limitations of existing video generation models, which often neglect audio and rely on cascaded pipelines, leading to high computational costs, error propagation, and audio-visual desynchronization. To overcome these challenges, we propose MOVA—the first open-source, end-to-end image-and-text-to-video-and-audio (IT2VA) generation model. Built upon a 32B-parameter Mixture-of-Experts architecture (with 18B activated per forward pass), MOVA supports LoRA-based fine-tuning, efficient inference, and prompt enhancement. It simultaneously generates semantically aligned, high-fidelity video and audio, including lip-synced speech, contextually appropriate sound effects, and background music. By releasing the model weights and a complete toolchain, this work aims to advance research in joint audio-visual synthesis.

multimodal modelingopen-sourcescalability

Hot Scholars

YW

Yao Wan

Huazhong University of Science and Technology
NLPProgramming LanguagesSoftware EngineeringLarge Language Models
HF

Hehe Fan

Zhejiang University
Deep learningComputer visionMultimediaAI for science
RZ

Renrui Zhang

Seed ByteDance & MMLab & PKU
Large Multimodal ModelGenerative ModelEmbodied AI
CL

Chenfei Liao

MPhil Student, HKUST(GZ)
Multi-modal PerceptionMLLMRobotic Vision
NC

Nan Cao

Professor, Intelligent Big Data Visualization Lab @ Tongji University
Visual AnalyticsInformation VisualizationVisualizationHuman-Computer Interaction