vision-language models

Designs, trains, and evaluates models and pipelines that jointly process visual (images, video) and textual inputs to produce aligned multimodal representations, build encoders/decoders and cross-modal interaction modules (e.g., cross-attention or contrastive objectives), and enable tasks such as image captioning, visual question answering, multimodal retrieval, and grounded language understanding. Work includes dataset curation and annotation, multimodal pretraining and fine‑tuning, and analysis of representation alignment, grounding, robustness, and inference for vision‑language systems.

vision-languagemodels

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-1.86
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$214K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Multimodal large language models (MLLMs) suffer from insufficient vision–language alignment, over-relying on linguistic priors while underutilizing fine-grained visual region understanding. This work presents the first systematic investigation into the internal visual comprehension mechanisms of MLLMs and proposes a novel paradigm—“visual depth enhancement” and “vision–language dynamic alignment”—to overcome the language-prior dominance bottleneck. Methodologically, we integrate attention mechanism analysis, visual feature disentanglement, cross-modal gated alignment, and token-level supervision targeting vision-dependent token prediction to strengthen visual representation learning and vision-guided language generation. Experiments demonstrate significant improvements: upstream vision-dependent token prediction accuracy increases notably, and average performance on vision-intensive tasks improves by 10 percentage points. These results validate that our paradigm effectively enhances multimodal alignment capability.

Enhancing visual comprehension in Multimodal Large Language ModelsImproving visual attention for better language generationReducing reliance on language priors in MLLMs

This paper addresses key challenges in multimodal large language models (MLLMs)—namely limited scalability, insufficient robustness, and difficulties in cross-modal alignment—within vision-language tasks. To tackle these, it proposes a comprehensive technical taxonomy covering the full MLLM stack and introduces a tripartite evaluation framework centered on scalability, robustness, and ethical alignment. Methodologically, the work conducts a systematic comparative analysis of core techniques, including Transformer-based fusion architectures, visual encoders (e.g., ViT, CLIP), cross-modal alignment learning, instruction tuning, and multi-stage training. It further synthesizes structured best practices for modeling across text, image, video, and audio modalities. The study delivers 12 representative application cases, empirically validating its framework and guidelines. Collectively, this work advances trustworthy deployment and cross-domain adoption of MLLMs through rigorous architectural analysis, principled evaluation, and actionable implementation insights.

Addressing challenges in scalability, robustness, and cross-modal learningExploring integration of text, images, video, and audio for cross-modal understandingSurveying multimodal large language models' architectures and applications

Multimodal Alignment and Fusion: A Survey

Nov 26, 2024
SL
Songtao Li
🏛️ Northeastern University | Peking University

This paper systematically analyzes four core challenges in multimodal alignment and fusion—cross-modal misalignment, semantic modality gap, computational bottlenecks, and data noise/heterogeneity—that hinder model robustness, generalizability, and scalability. To address them, the authors propose a unified taxonomy covering 200+ studies, revealing an integrated “alignment–fusion” co-design paradigm. They further introduce three principled solution pathways: (1) noise-robust learning, (2) heterogeneous representation modeling, and (3) few-shot cross-modal transfer—unifying contrastive learning, cross-modal attention, latent-space alignment, graph neural networks, and differentiable architecture search. Extensive experiments on social media analysis, medical imaging, and sentiment recognition validate the efficacy of the framework. The work distills actionable design principles for scalable, robust, and generalizable multimodal learning, establishing a new benchmark for both theoretical research and industrial deployment.

Addressing cross-modal misalignment and computational bottlenecks challengesExploring applications from social media to medical imagingSurveying multimodal alignment and fusion techniques in machine learning

Everything is a Video: Unifying Modalities through Next-Frame Prediction

Nov 15, 2024
GT
G. Thomas
🏛️ Durham University

Traditional multimodal approaches rely on modality-specific encoders and late fusion, limiting scalability and cross-modal generalization. This work proposes a unified paradigm that reformulates diverse multimodal tasks—including text, image, audio, and video processing—as “next-frame prediction,” with all inputs and outputs represented as serialized video frames, enabling end-to-end, single-model inference. It pioneers the task-redefinition strategy—previously confined to NLP—within multimodal learning, eliminating modality-specific design in favor of a fully modality-agnostic architecture. The framework integrates cross-modal tokenization, sequential frame representation, autoregressive modeling, and a shared Transformer decoder, supporting joint multi-task pretraining. Empirical evaluation demonstrates strong zero-shot and few-shot generalization across text-to-text, image-to-text, video-to-video, video-to-text, and audio-to-text tasks. Crucially, adaptation requires only lightweight task-specific heads, underscoring its flexibility and efficiency.

Enabling seamless knowledge transfer across different tasks and modalitiesOvercoming modality-specific encoder limitations in multimodal learningUnifying diverse modalities via next-frame prediction

EMMA: Efficient Visual Alignment in Multi-Modal LLMs

Oct 02, 2024
SG
Sara Ghazanfari
🏛️ New York University

To address inefficient vision-language fusion, reliance on complex adapter modules, and large-scale training data in multimodal large language models (MLLMs), this paper proposes EMMA—a lightweight cross-modal alignment module. Methodologically, EMMA introduces: (1) an efficient early-fusion mechanism with <0.2% parameter overhead, leveraging instruction-conditioned visual feature reweighting and a lightweight cross-attention adapter to generate instruction-aware visual representations; (2) an interpretable analytical framework that elucidates the intrinsic mechanisms of cross-modal alignment; and (3) comprehensive evaluation demonstrating an average 9.3% improvement across diverse domain-specific and general-purpose benchmarks. EMMA significantly mitigates hallucination and enhances robustness while preserving model simplicity and computational efficiency. The approach achieves substantial performance gains without architectural bloat or extensive retraining, offering a principled trade-off between effectiveness and parsimony in MLLM design.

Efficient fusion of visual and textual encodings in MLLMsImproving task-specific adaptability and performanceReducing model complexity and training data needs

Latest Papers

What's happening recently
View more

This work explores the design space of natively multimodal foundation models, addressing how to effectively integrate vision and language beyond conventional language modeling. Building upon the Transfusion framework, the authors propose a unified pretraining approach from scratch that jointly leverages next-token prediction and diffusion-based generation, augmented with a Representation Autoencoder (RAE) to unify visual representations for both understanding and generation. The study reveals the complementary nature and asymmetric scaling behavior of vision and language data—where vision benefits more substantially from increased data volume—and employs a Mixture-of-Experts (MoE) architecture to enable efficient modality specialization and model expansion. Experiments demonstrate that unified pretraining naturally induces world modeling capabilities and significantly enhances performance on downstream tasks, laying a foundation for truly integrated multimodal foundation models.

foundation modelsmodality asymmetrymultimodal pretraining

Remodeling Semantic Relationships in Vision-Language Fine-Tuning

Nov 11, 2025
XW
Xiangyang Wu
🏛️ Hangzhou International Innovation Institute, Beihang University | Beihang University | Nanyang Technological University | University of Leicester

Existing vision-language fine-tuning methods often neglect image-internal semantic relationships emphasized by text, leading to suboptimal cross-modal alignment. To address this, we propose a semantics- and relation-driven multimodal alignment framework. First, a multi-level visual encoder explicitly models fine-grained intra-image semantic relations. Second, we design a transferable cross-attention mechanism that dynamically filters low-relevance vision–text feature pairs at the global level, enabling robust multimodal fusion. Third, semantic grouping projection is introduced to enhance cross-modal interaction. Our framework demonstrates strong generalizability, validated across eight mainstream foundation models. It achieves significant improvements over state-of-the-art methods on both visual question answering and image captioning benchmarks. Empirical results underscore the critical role of explicit intra-image relational modeling in enhancing the quality of cross-modal alignment.

Addressing overlooked semantic relationships in vision-language fine-tuningEnhancing visual question answering and image captioning performanceImproving multimodal alignment through semantic relationship modeling

Real-world multimodal video data suffers from high acquisition costs and limited diversity, hindering the training of large-scale multitask video understanding models. To address this challenge, this work proposes the first unified synthetic data generation framework capable of automatically producing unlimited, multitask-compatible multimodal video data. The approach introduces a visual question answering (VQA)-based fine-tuning strategy that replaces conventional caption- or instruction-based supervision with structured question-answer pairs to enhance the model’s visual reasoning and localization capabilities. Remarkably, models trained exclusively on this synthetic data achieve performance on par with or even surpassing fully supervised baselines on three distinct tasks—video object counting, video question answering, and video segmentation—demonstrating strong generalization and effectiveness across real-world benchmarks.

data annotationlarge language modelsmultimodal video understanding

This work systematically investigates the core mechanisms of modality interaction in multimodal pretraining, focusing on knowledge transfer, synergistic effects, and fusion timing. Through experiments on both synthetic and large-scale real-world datasets, it provides the first empirical evidence of asymmetric cross-modal knowledge flow and demonstrates that data complexity governs whether modalities exhibit synergy or competition. The study further validates that early unified fusion consistently outperforms late alignment. Leveraging an architecture featuring shared attention and normalization layers with modality-specific feedforward components, the proposed approach is evaluated on a 13.5B mixture-of-experts model trained on 2 trillion tokens, confirming its effectiveness. Additionally, the paper introduces a highly efficient pretraining strategy that achieves strong generative performance using only 5% of the typical computational budget.

early unificationknowledge flowmodality interaction

CACARA: Cross-Modal Alignment Leveraging a Text-Centric Approach for Cost-Effective Multimodal and Multilingual Learning

Nov 29, 2025
DA
Diego A. B. Moreira
🏛️ Universidade Estadual de Campinas | Universidade Federal de Goiás | University of Sheffield

To address the bottleneck where cross-modal alignment and multilingual capability expansion traditionally require large-scale multimodal/multilingual pretraining, this paper proposes CACARA—a text-centric cross-modal alignment architecture. Its core innovation lies in enabling emergent audio–text retrieval capabilities across 100 languages by fine-tuning only the newly introduced modality encoder on English-aligned data, while keeping the pretrained text encoder frozen. CACARA integrates parameter-efficient fine-tuning with a monolingual-to-multilingual transfer mechanism, achieving low-cost capability extension without compromising original knowledge. Extensive experiments demonstrate that CACARA achieves up to a 14.24-percentage-point improvement in Recall@1 on audio–text retrieval tasks, outperforming state-of-the-art multimodal models, while maintaining training costs comparable to monolingual baselines.

Enables multilingual support from monolingual training dataIntegrates new modalities without full model retrainingReduces computational cost while improving retrieval performance

Hot Scholars

DT

Dzmitry Tsetserukou

Associate Professor, Skolkovo Institute of Science and Technology (Skoltech)
RoboticsHapticsUAV SwarmAI
XX

Xiangyang Xue

Professor of Computer Science, Fudan University
Computer VisionPattern RecognitionMachine Learning
MA

Miguel Altamirano Cabrera

Research Scientist, Skolkovo Institute of Science and Technology
HapticsRoboticsTactile SensationComputer Vision
SD

Shuangrui Ding

The Chinese University of Hong Kong
Computer Vision
YZ

Yuhang Zang

Shanghai AI Laboratory
Natural Language ProcessingVision Language Model