Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation

๐Ÿ“… 2025-05-29
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This work investigates whether understanding and generation tasks in unified vision-language models (VLMs) can mutually enhance and generalize across modalities. To this end, we construct a real-world scenario-aligned multimodal dataset and systematically evaluate bidirectional transfer capabilities of diverse unified architectures under mixed-task training, complemented by quantitative analysis and ablation studies. Our key contributions are threefold: (1) We provide the first empirical evidence that knowledge acquired from generative tasks effectively transfers to discriminative understanding tasksโ€”crucially, this transfer occurs within the base language model itself, not merely through modality adapters; (2) We identify alignment quality in the input-output multimodal embedding space as a critical determinant of cross-task generalization; and (3) Mixed-task training substantially improves bidirectional generalization performance, with gains scaling favorably with data volume. These findings offer pivotal empirical support for the architectural necessity of unified VLMs.

Technology Category

Machine Learning: Multimodal LearningComputer Vision: Multi-modal VisionNatural Language Processing: Language Grounding & Multi-modal NLP

Application Category

Semantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsSearch and Retrieval-Augmented AI: Retrieval-Augmented Generation (RAG) and multi-modal RAGEconomics, Online Markets and Human Computation: Uses of LLMs and GenAI for marketplace design, bidding, and strategic interactions
๐Ÿ“ Abstract
Recent advancements in unified vision-language models (VLMs), which integrate both visual understanding and generation capabilities, have attracted significant attention. The underlying hypothesis is that a unified architecture with mixed training on both understanding and generation tasks can enable mutual enhancement between understanding and generation. However, this hypothesis remains underexplored in prior works on unified VLMs. To address this gap, this paper systematically investigates the generalization across understanding and generation tasks in unified VLMs. Specifically, we design a dataset closely aligned with real-world scenarios to facilitate extensive experiments and quantitative evaluations. We evaluate multiple unified VLM architectures to validate our findings. Our key findings are as follows. First, unified VLMs trained with mixed data exhibit mutual benefits in understanding and generation tasks across various architectures, and this mutual benefits can scale up with increased data. Second, better alignment between multimodal input and output spaces will lead to better generalization. Third, the knowledge acquired during generation tasks can transfer to understanding tasks, and this cross-task generalization occurs within the base language model, beyond modality adapters. Our findings underscore the critical necessity of unifying understanding and generation in VLMs, offering valuable insights for the design and optimization of unified VLMs.
Problem

Research questions and friction points this paper is trying to address.

Investigates generalization between vision-language understanding and generation tasks
Examines mutual benefits of mixed training in unified VLMs
Explores cross-task knowledge transfer within multimodal architectures
Innovation

Methods, ideas, or system contributions that make the work stand out.

Unified VLMs enhance understanding and generation tasks
Better multimodal alignment improves generalization performance
Generation task knowledge transfers to understanding tasks
๐Ÿ’ผ Related Jobs
No related jobs found.
J
Jihai Zhang
The Chinese University of Hong Kong
T
Tianle Li
The Chinese University of Hong Kong
Linjie Li
Linjie Li
Microsoft
Vision and Language
Zhengyuan Yang
Zhengyuan Yang
Principal Researcher, Microsoft
Computer VisionMultimediaMultimodalPost-TrainingAgentic RL
Y
Yu Cheng
The Chinese University of Hong Kong