Score
Designs and implements unified vision–language models that jointly process visual and textual inputs to produce multiple structured outputs (for example classification, detection, and segmentation) from a single architecture, with an emphasis on interpretation-first or explanation-capable outputs. Develops the multitask heads, training and inference procedures that enable operation without external labels at test time and support human-in-the-loop correction and interactive interpretation.
Large language models (LLMs) inherently lack native visual understanding, limiting their applicability in multimodal scenarios. This work presents a systematic survey of vision-language models (VLMs), introducing— for the first time—a three-tier taxonomy grounded in modal input/output capabilities: comprehension-only, generation-only, and full-modality VLMs. We unify analysis across architectural design, training data composition, robustness properties, and benchmark performance (e.g., VQAv2, COCO Caption). Through comprehensive literature review, architectural decomposition, and cross-benchmark evaluation, we analyze over 100 state-of-the-art works to construct a technology evolution map. Our key contributions are: (1) the first scalable, capability-aware classification and evaluation framework for VLMs; (2) a precise delineation of current performance boundaries; and (3) identification of three critical future directions—embodied intelligence, robust multimodal reasoning, and efficient scaling—establishing an authoritative reference for the VLM research community.
Existing surveys often treat vision and language modalities in isolation, lacking a unified perspective on the evolution of perceptual capabilities in Multimodal Large Language Models (MLLMs). This work presents the first systematic review framed around human-like innate visio-linguistic joint perception, introducing a five-stage taxonomy that captures paradigmatic shifts and emphasizes cross-modal synergy over modality separation. Through comprehensive literature synthesis, paradigm categorization, and cross-modal analysis, the study systematically examines representative architectures, training strategies, and key milestones, tracing the trajectory of multimodal perception from structural fusion toward collaborative understanding. It further identifies current challenges and outlines a clear research roadmap toward Artificial General Intelligence.
This work addresses the parameter efficiency challenge in visual–language large models (VLLMs) for effective cross-modal fusion. We systematically analyze 34 state-of-the-art VLLMs and, for the first time, unify their training paradigms into three categories—single-stage fine-tuning, two-stage fine-tuning, and direct adaptation—establishing the first taxonomy of VLLM efficiency grounded in training methodology. Our study fills a critical gap by providing the first systematic analysis of direct adaptation, empirically demonstrating that it achieves over 90% of two-stage fine-tuning performance with less than 1% parameter overhead. We comprehensively examine core components—including LLM backbones, vision encoders, multimodal fusion architectures, parameter-efficient adaptation techniques (e.g., LoRA, Adapters), and evaluation protocols—and synthesize key benchmarks and metrics. The work delivers both a theoretical framework and empirical evidence to advance efficient multimodal modeling.
This work investigates whether understanding and generation tasks in unified vision-language models (VLMs) can mutually enhance and generalize across modalities. To this end, we construct a real-world scenario-aligned multimodal dataset and systematically evaluate bidirectional transfer capabilities of diverse unified architectures under mixed-task training, complemented by quantitative analysis and ablation studies. Our key contributions are threefold: (1) We provide the first empirical evidence that knowledge acquired from generative tasks effectively transfers to discriminative understanding tasks—crucially, this transfer occurs within the base language model itself, not merely through modality adapters; (2) We identify alignment quality in the input-output multimodal embedding space as a critical determinant of cross-task generalization; and (3) Mixed-task training substantially improves bidirectional generalization performance, with gains scaling favorably with data volume. These findings offer pivotal empirical support for the architectural necessity of unified VLMs.
This work addresses the prevalent text-dominant bias in existing vision-language models (VLMs), where visual signals are treated merely as passive inputs, leading to the loss of fine-grained visual details and coarse-grained multimodal understanding. To overcome this limitation, we propose Youtu-VL, a novel framework that introduces the Vision-Language Unified Autoregressive Supervision (VLUAS) paradigm. VLUAS unifies visual and linguistic tokens into a single autoregressive prediction sequence, enabling visual tokens to serve as prediction targets rather than just contextual inputs. This approach breaks away from conventional text-centric training paradigms and supports a wide range of vision-centric tasks without task-specific customization. Extensive experiments demonstrate that Youtu-VL achieves competitive performance on both general multimodal benchmarks and vision-intensive tasks, significantly enhancing visual detail preservation and joint multimodal modeling capabilities.
How do language models—lacking explicit visual pretraining—achieve image understanding? Method: We systematically analyze 16 multimodal large language models (MLLMs) spanning four architectural families and four parameter scales. Introducing the concept of “vision-preferring attention heads,” we identify such heads via attention behavior analysis, statistical modeling of attention weights, and cross-scale ablation experiments, empirically validating their strong, consistent focus on visual tokens. Contribution: We are the first to discover and formally define this generalizable, modular visual-perception substructure within LLMs. Our work reveals the pivotal role of attention mechanisms in cross-modal adaptation, demonstrating how vision-preferring heads mediate text–vision alignment. This provides an interpretable, spatially localizable mechanism underlying joint text–vision representation learning, thereby advancing research toward transparent, controllable, and analyzable multimodal foundation models.
This work investigates how visual tokens are processed within the language model component of vision-language models (VLMs). Focusing on the LLaVA architecture, we propose a three-tier interpretability framework: (1) Through hierarchical token attribution and representation alignment visualization, we first observe that visual token representations progressively align with the textual vocabulary space across layers; (2) We identify critical semantic information concentrated in the final visual token, revealing a fact-retrieval–like predictive mechanism analogous to pure language models; (3) Systematic ablation experiments demonstrate that removing object-specific visual tokens degrades recognition accuracy by over 70%, while representational interpretability markedly increases with network depth. Our study establishes the first fine-grained, cross-layer, token-level interpretability paradigm for vision–language fusion in multimodal models, offering novel mechanistic insights into how visual information is encoded, aligned, and utilized within language decoders.
This work proposes UniMRG, a unified multimodal framework that addresses the limited synergy between visual understanding and generation in existing models. By integrating auxiliary generative tasks—such as pixel reconstruction, depth estimation, and semantic segmentation—within a single architecture, UniMRG enables bidirectional enhancement between comprehension and synthesis. The method employs an architecture-agnostic post-training strategy, uniquely leveraging multitask generation to retroactively improve visual understanding capabilities. Experimental results demonstrate that UniMRG significantly advances performance in fine-grained perception, spatial relationship modeling, and hallucination suppression, while simultaneously enhancing generation quality. These findings validate the efficacy of the proposed understanding-generation co-evolution mechanism within a unified model.
This work addresses the challenge that existing vision models typically rely on task-specific architectures for diverse structured prediction tasks—such as segmentation, depth estimation, and pose generation—hindering a unified approach. To overcome this limitation, the authors propose RINO, a novel framework that, for the first time, uniformly encodes various structured visual inputs and outputs (e.g., masks, depth maps, keypoints) into RGB images, thereby reformulating a wide range of vision tasks as a general RGB-to-RGB image editing problem. Built upon a shared encoder-decoder backbone, RINO establishes a language-model-like universal visual interface capable of zero-shot cross-task transfer without requiring task-specific fine-tuning. Experiments demonstrate that RINO achieves competitive zero-shot performance on both dense understanding and conditional generation tasks.
This work proposes a unified vision model that transcends task-specific architectures traditionally required in computer vision by formulating diverse visual tasks—including detection, segmentation, and geometric prediction—as multimodal generation problems. The model is driven solely by natural language instructions (optionally augmented with visual prompts) to produce text, images, or hybrid outputs from a single architecture, eliminating the need for specialized heads or modular designs. It introduces the SenseNova-Vision Corpus, a large-scale dataset of vision-language instruction-response pairs, and leverages off-the-shelf pretrained multimodal models refined through instruction tuning and joint multimodal generation strategies. This end-to-end framework supports compositional, language-defined tasks and achieves performance on par with or superior to specialized systems across structured understanding, dense geometric prediction, segmentation, and multi-view geometry benchmarks.
Existing evaluation methods treat visual generation and understanding as disjoint capabilities, failing to holistically assess the system-level performance of unified multimodal models. This work proposes a Self-Generated Understanding (SGU) framework, introducing a novel annotation-free semantic closed-loop evaluation paradigm: a model first describes an image, then reconstructs visual content from its own generated text, and finally performs zero-shot reasoning on the reconstructed output. SGU uniquely integrates generation and understanding into a single assessment pipeline, leveraging a self-feedback mechanism to uncover latent deficiencies that manifest specifically within the model’s self-generated context. Experiments reveal that even high-performing models exhibit substantially degraded reasoning capabilities under SGU, exposing systemic limitations invisible to conventional isolated evaluations.
Current evaluations of unified multimodal models often treat understanding and generation capabilities in isolation, neglecting their synergistic interplay. This work introduces Unison, a benchmark comprising 2,169 high-quality task samples, which for the first time systematically assesses model synergy along three dimensions: internal consistency, mutual guidance, and reciprocal enhancement. To enable fine-grained analysis, the authors design both unified and decoupled diagnostic pathways and develop Unison-Judge, an automatic scoring model aligned with human preferences. Their findings uncover critical limitations in existing models’ ability to jointly perform understanding and generation, offering clear directions for future research. The dataset and evaluation toolkit are publicly released to support further advancements in multimodal foundation models.