Score
Designs, trains, and evaluates models and pipelines that jointly process visual (images, video) and textual inputs to produce aligned multimodal representations, build encoders/decoders and cross-modal interaction modules (e.g., cross-attention or contrastive objectives), and enable tasks such as image captioning, visual question answering, multimodal retrieval, and grounded language understanding. Work includes dataset curation and annotation, multimodal pretraining and fine‑tuning, and analysis of representation alignment, grounding, robustness, and inference for vision‑language systems.
Large language models (LLMs) inherently lack native visual understanding, limiting their applicability in multimodal scenarios. This work presents a systematic survey of vision-language models (VLMs), introducing— for the first time—a three-tier taxonomy grounded in modal input/output capabilities: comprehension-only, generation-only, and full-modality VLMs. We unify analysis across architectural design, training data composition, robustness properties, and benchmark performance (e.g., VQAv2, COCO Caption). Through comprehensive literature review, architectural decomposition, and cross-benchmark evaluation, we analyze over 100 state-of-the-art works to construct a technology evolution map. Our key contributions are: (1) the first scalable, capability-aware classification and evaluation framework for VLMs; (2) a precise delineation of current performance boundaries; and (3) identification of three critical future directions—embodied intelligence, robust multimodal reasoning, and efficient scaling—establishing an authoritative reference for the VLM research community.
Multimodal large language models (MLLMs) suffer from insufficient vision–language alignment, over-relying on linguistic priors while underutilizing fine-grained visual region understanding. This work presents the first systematic investigation into the internal visual comprehension mechanisms of MLLMs and proposes a novel paradigm—“visual depth enhancement” and “vision–language dynamic alignment”—to overcome the language-prior dominance bottleneck. Methodologically, we integrate attention mechanism analysis, visual feature disentanglement, cross-modal gated alignment, and token-level supervision targeting vision-dependent token prediction to strengthen visual representation learning and vision-guided language generation. Experiments demonstrate significant improvements: upstream vision-dependent token prediction accuracy increases notably, and average performance on vision-intensive tasks improves by 10 percentage points. These results validate that our paradigm effectively enhances multimodal alignment capability.
This paper addresses key challenges in multimodal large language models (MLLMs)—namely limited scalability, insufficient robustness, and difficulties in cross-modal alignment—within vision-language tasks. To tackle these, it proposes a comprehensive technical taxonomy covering the full MLLM stack and introduces a tripartite evaluation framework centered on scalability, robustness, and ethical alignment. Methodologically, the work conducts a systematic comparative analysis of core techniques, including Transformer-based fusion architectures, visual encoders (e.g., ViT, CLIP), cross-modal alignment learning, instruction tuning, and multi-stage training. It further synthesizes structured best practices for modeling across text, image, video, and audio modalities. The study delivers 12 representative application cases, empirically validating its framework and guidelines. Collectively, this work advances trustworthy deployment and cross-domain adoption of MLLMs through rigorous architectural analysis, principled evaluation, and actionable implementation insights.
This paper systematically analyzes four core challenges in multimodal alignment and fusion—cross-modal misalignment, semantic modality gap, computational bottlenecks, and data noise/heterogeneity—that hinder model robustness, generalizability, and scalability. To address them, the authors propose a unified taxonomy covering 200+ studies, revealing an integrated “alignment–fusion” co-design paradigm. They further introduce three principled solution pathways: (1) noise-robust learning, (2) heterogeneous representation modeling, and (3) few-shot cross-modal transfer—unifying contrastive learning, cross-modal attention, latent-space alignment, graph neural networks, and differentiable architecture search. Extensive experiments on social media analysis, medical imaging, and sentiment recognition validate the efficacy of the framework. The work distills actionable design principles for scalable, robust, and generalizable multimodal learning, establishing a new benchmark for both theoretical research and industrial deployment.
Traditional multimodal approaches rely on modality-specific encoders and late fusion, limiting scalability and cross-modal generalization. This work proposes a unified paradigm that reformulates diverse multimodal tasks—including text, image, audio, and video processing—as “next-frame prediction,” with all inputs and outputs represented as serialized video frames, enabling end-to-end, single-model inference. It pioneers the task-redefinition strategy—previously confined to NLP—within multimodal learning, eliminating modality-specific design in favor of a fully modality-agnostic architecture. The framework integrates cross-modal tokenization, sequential frame representation, autoregressive modeling, and a shared Transformer decoder, supporting joint multi-task pretraining. Empirical evaluation demonstrates strong zero-shot and few-shot generalization across text-to-text, image-to-text, video-to-video, video-to-text, and audio-to-text tasks. Crucially, adaptation requires only lightweight task-specific heads, underscoring its flexibility and efficiency.
To address inefficient vision-language fusion, reliance on complex adapter modules, and large-scale training data in multimodal large language models (MLLMs), this paper proposes EMMA—a lightweight cross-modal alignment module. Methodologically, EMMA introduces: (1) an efficient early-fusion mechanism with <0.2% parameter overhead, leveraging instruction-conditioned visual feature reweighting and a lightweight cross-attention adapter to generate instruction-aware visual representations; (2) an interpretable analytical framework that elucidates the intrinsic mechanisms of cross-modal alignment; and (3) comprehensive evaluation demonstrating an average 9.3% improvement across diverse domain-specific and general-purpose benchmarks. EMMA significantly mitigates hallucination and enhances robustness while preserving model simplicity and computational efficiency. The approach achieves substantial performance gains without architectural bloat or extensive retraining, offering a principled trade-off between effectiveness and parsimony in MLLM design.
This work explores the design space of natively multimodal foundation models, addressing how to effectively integrate vision and language beyond conventional language modeling. Building upon the Transfusion framework, the authors propose a unified pretraining approach from scratch that jointly leverages next-token prediction and diffusion-based generation, augmented with a Representation Autoencoder (RAE) to unify visual representations for both understanding and generation. The study reveals the complementary nature and asymmetric scaling behavior of vision and language data—where vision benefits more substantially from increased data volume—and employs a Mixture-of-Experts (MoE) architecture to enable efficient modality specialization and model expansion. Experiments demonstrate that unified pretraining naturally induces world modeling capabilities and significantly enhances performance on downstream tasks, laying a foundation for truly integrated multimodal foundation models.
Existing vision-language fine-tuning methods often neglect image-internal semantic relationships emphasized by text, leading to suboptimal cross-modal alignment. To address this, we propose a semantics- and relation-driven multimodal alignment framework. First, a multi-level visual encoder explicitly models fine-grained intra-image semantic relations. Second, we design a transferable cross-attention mechanism that dynamically filters low-relevance vision–text feature pairs at the global level, enabling robust multimodal fusion. Third, semantic grouping projection is introduced to enhance cross-modal interaction. Our framework demonstrates strong generalizability, validated across eight mainstream foundation models. It achieves significant improvements over state-of-the-art methods on both visual question answering and image captioning benchmarks. Empirical results underscore the critical role of explicit intra-image relational modeling in enhancing the quality of cross-modal alignment.
Real-world multimodal video data suffers from high acquisition costs and limited diversity, hindering the training of large-scale multitask video understanding models. To address this challenge, this work proposes the first unified synthetic data generation framework capable of automatically producing unlimited, multitask-compatible multimodal video data. The approach introduces a visual question answering (VQA)-based fine-tuning strategy that replaces conventional caption- or instruction-based supervision with structured question-answer pairs to enhance the model’s visual reasoning and localization capabilities. Remarkably, models trained exclusively on this synthetic data achieve performance on par with or even surpassing fully supervised baselines on three distinct tasks—video object counting, video question answering, and video segmentation—demonstrating strong generalization and effectiveness across real-world benchmarks.
This work systematically investigates the core mechanisms of modality interaction in multimodal pretraining, focusing on knowledge transfer, synergistic effects, and fusion timing. Through experiments on both synthetic and large-scale real-world datasets, it provides the first empirical evidence of asymmetric cross-modal knowledge flow and demonstrates that data complexity governs whether modalities exhibit synergy or competition. The study further validates that early unified fusion consistently outperforms late alignment. Leveraging an architecture featuring shared attention and normalization layers with modality-specific feedforward components, the proposed approach is evaluated on a 13.5B mixture-of-experts model trained on 2 trillion tokens, confirming its effectiveness. Additionally, the paper introduces a highly efficient pretraining strategy that achieves strong generative performance using only 5% of the typical computational budget.
To address the bottleneck where cross-modal alignment and multilingual capability expansion traditionally require large-scale multimodal/multilingual pretraining, this paper proposes CACARA—a text-centric cross-modal alignment architecture. Its core innovation lies in enabling emergent audio–text retrieval capabilities across 100 languages by fine-tuning only the newly introduced modality encoder on English-aligned data, while keeping the pretrained text encoder frozen. CACARA integrates parameter-efficient fine-tuning with a monolingual-to-multilingual transfer mechanism, achieving low-cost capability extension without compromising original knowledge. Extensive experiments demonstrate that CACARA achieves up to a 14.24-percentage-point improvement in Recall@1 on audio–text retrieval tasks, outperforming state-of-the-art multimodal models, while maintaining training costs comparable to monolingual baselines.