multimodal large language models

Designs, builds, and evaluates large language model systems that jointly process and generate multiple data modalities (e.g., text plus images, audio, or video), including multimodal encoders/decoders, cross‑modal fusion and alignment modules, and inference pipelines. Creates and curates multimodal training and evaluation datasets, implements fine‑tuning and scaling workflows, and analyzes model performance, robustness, and cross‑modal behavior for integrated multimodal applications.

multimodallargelanguagemodels

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-3.29
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$214K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Unifying Multimodal Large Language Model Capabilities and Modalities via Model Merging

May 26, 2025
YW
Yongxian Wei
🏛️ Tsinghua University | Huawei | Sun Yat-sen University | Nanyang Technological University

A standardized, integrated benchmark for training and evaluating multimodal large language models (MLLMs) is lacking, hindering systematic research on cross-modal model merging. Method: We introduce the first dedicated MLLM model merging benchmark, covering diverse tasks—including visual question answering, geometric reasoning, chart understanding, OCR, and localization—and supporting both LoRA and full-parameter fine-tuned models. We systematically investigate vision–language, audio–language, and video–language fusion pathways. Furthermore, we propose a novel fusion method integrating task vector denoising, interactive loss optimization, LoRA adapter fusion, and multimodal alignment. Results: Our approach achieves an average performance gain of 2.48% across benchmarks, surpassing unimodal expert models without additional training data. This constitutes the first empirical demonstration that multimodal complementarity can yield stronger, general-purpose Omni-language models.

Improving model merging algorithms for better performanceLack of benchmark for merging multimodal large language modelsNeed to combine different modalities into unified model

This paper addresses key challenges in multimodal large language models (MLLMs)—namely limited scalability, insufficient robustness, and difficulties in cross-modal alignment—within vision-language tasks. To tackle these, it proposes a comprehensive technical taxonomy covering the full MLLM stack and introduces a tripartite evaluation framework centered on scalability, robustness, and ethical alignment. Methodologically, the work conducts a systematic comparative analysis of core techniques, including Transformer-based fusion architectures, visual encoders (e.g., ViT, CLIP), cross-modal alignment learning, instruction tuning, and multi-stage training. It further synthesizes structured best practices for modeling across text, image, video, and audio modalities. The study delivers 12 representative application cases, empirically validating its framework and guidelines. Collectively, this work advances trustworthy deployment and cross-domain adoption of MLLMs through rigorous architectural analysis, principled evaluation, and actionable implementation insights.

Addressing challenges in scalability, robustness, and cross-modal learningExploring integration of text, images, video, and audio for cross-modal understandingSurveying multimodal large language models' architectures and applications

Towards Efficient Large Multimodal Model Serving

Feb 02, 2025
HQ
Haoran Qiu
🏛️ Microsoft | University of Virginia

Multimodal large language models (LMMs) suffer from unstable inference performance, high resource consumption, and severe cross-modal request interference during production deployment. Method: We systematically analyze multi-stage inference behavior and resource contention patterns across six open-source LMMs, comparing decoder-only and cross-attention architectures. We propose a decoupled serving architecture enabling per-stage independent resource allocation and elastic scaling, and introduce stage co-location optimization—jointly maximizing throughput and resource utilization under latency constraints. Our approach integrates multimodal request trajectory modeling, heterogeneous resource scheduling, and compute-memory co-optimization. Contribution/Results: Evaluation shows up to 2.3× higher throughput, 41% reduction in tail latency, and 37% improvement in resource utilization compared to state-of-the-art baselines.

Inference PerformanceMultimodal AI ModelsResource Consumption

Towards LLM-Centric Multimodal Fusion: A Survey on Integration Strategies and Techniques

Jun 05, 2025
JA
Jisu An
🏛️ Seoul National University | University of California San Diego | Chung-Ang University

Existing research lacks a systematic analysis of modality–language backbone integration mechanisms in multimodal large language models (MLLMs). Method: We systematically survey 125 MLLMs published between 2021 and 2025, proposing the first large-language-model-centric three-dimensional taxonomy—architectural integration, representation learning, and training paradigms—to unify cross-modal alignment pathways. Leveraging bibliometric analysis, architectural decoupling, and fine-grained modeling of embeddings and loss functions, we characterize evolutionary patterns in fusion granularity, joint representation design, and objective function development. Contribution/Results: This work fills a critical theoretical gap in structured analysis of MLLM fusion mechanisms, identifies shared bottlenecks—including misaligned modality-specific feature hierarchies and suboptimal cross-modal optimization objectives—and delivers a reusable theoretical framework and practical guidelines for developing robust, scalable multimodal foundation models.

Analysis of 125 MLLMs to identify emerging patternsClassification framework for MLLMs based on key dimensionsSystematic understanding of multimodal integration with LLMs

OmniBench: Towards The Future of Universal Omni-Language Models

Sep 23, 2024
YL
Yizhi Li
🏛️ University of Manchester | 01.ai | Queen Mary University of London | Hongkong University of Science and Technology | Nanjing University | Dartmouth College

Existing open-source multimodal large language models (MLLMs) exhibit significant deficiencies in joint visual-auditory-textual understanding and reasoning, achieving only ~50% instruction-following accuracy on trilingual multimodal tasks. Method: We introduce OmniBench—the first benchmark for trilingual multimodal collaborative reasoning—and formalize the omni-language model (OLM), a unified architecture capable of jointly processing visual, auditory, and textual (V-A-T) inputs. We construct OmniBench via expert human annotation across diverse trilingual multimodal tasks and curate OmniInstruct, a large-scale instruction-tuning dataset comprising 96K samples. Our methodology integrates cross-modal alignment modeling, trilingual multimodal instruction tuning, and a human-in-the-loop evaluation framework. Contribution/Results: Experiments reveal severe generalization limitations of current open-source OLMs on trilingual multimodal tasks; OmniInstruct substantially improves their reasoning performance. This work establishes a novel evaluation paradigm, provides high-quality resources, and outlines a technical pathway for advancing trilingual multimodal foundation models.

Addressing limitations in instruction-following and reasoning in tri-modal contextsDeveloping robust tri-modal integration techniques for omni-language modelsEvaluating models' ability to process visual, acoustic, and textual inputs simultaneously

Latest Papers

What's happening recently
View more

Existing model-serving frameworks struggle to efficiently support increasingly complex composite multimodal models. To address this challenge, this work proposes M*, a novel system that introduces a modular Walk Graph abstraction to uniformly represent composite AI models as dataflow graphs. This abstraction enables flexible composition of arbitrary model components, cluster deployment, and model-agnostic distributed runtime optimizations. By leveraging a graph traversal mechanism, M* efficiently handles cross-modal, multitask requests, achieving significant performance gains: it reduces end-to-end latency by 20% over vLLM-Omni in text-to-image generation, improves real-time factor by 2.9× and throughput by 2.7× in text-to-speech tasks, and accelerates robot planning workloads by up to 12.5×.

architectural diversitycomposite architecturesmodel serving

This work addresses the limitations of existing approaches that predominantly rely on non-native late-fusion architectures and lack a systematic definition or unified framework for native multimodal modeling. The paper proposes a formal roadmap for Native Multimodal Modeling (NMM), introducing, for the first time, a rigorous formulation of “architectural nativeness.” Leveraging an input–output duality perspective, it categorizes models into three types: Multi-to-Text, Multi-to-Target, and Multi-to-Multi. A full-stack, industrial-grade NMM framework is developed, encompassing data governance, early/mid-stage fusion, end-to-end training, and deployment, all realized through a unified Transformer architecture that enables symbiotic cross-modal understanding and generation. This paradigm demonstrates superior performance and strong scalability across multiple tasks, offering a clear pathway toward truly native multimodal models.

architectural nativitymultimodal architecturemultimodal fusion

This study addresses the unclear trade-offs among diverse capabilities—such as general visual understanding, structured reasoning, and fine-grained OCR—in multimodal instruction tuning under mixed data regimes, particularly the lack of systematic investigation into how data organization influences these trade-offs. Treating data scheduling as a first-order design variable while holding model architecture and optimization settings fixed, the work compares four strategies: direct mixing, curriculum learning, balanced sampling, and reverse curriculum. Results demonstrate that curriculum-based training—sequencing tasks from general comprehension to specialized skills—achieves superior overall performance and structured reasoning while accelerating convergence. Balanced sampling improves OCR accuracy at the cost of capability imbalance, whereas reverse curriculum degrades performance and induces optimization instability. The findings highlight the critical role of training sequence in shaping the capability distribution of multimodal models.

capability trade-offsdata organizationmultimodal instruction tuning

Multimodal Large Language Models (MLLMs) combine the natural language understanding and generation capabilities of LLMs with perception skills in modalities such as image and audio, representing a key advancement in contemporary AI. This chapter presents the main fundamentals of MLLMs and emblematic models. Practical techniques for preprocessing, prompt engineering, and building multimodal pipelines with LangChain and LangGraph are also explored. For further practical study, supplementary material is publicly available online: https://github.com/neemiasbsilva/MLLMs-Teoria-e-Pratica. Finally, the chapter discusses the challenges and highlights promising trends.

AI integrationMultimodal Large Language Modelsmultimodal perception

This work addresses the high cost and combinatorial explosion associated with optimizing data mixing ratios in cross-domain supervised fine-tuning of multimodal large language models. The authors propose an innovative approach that constructs a surrogate model by linearly merging parameters from domain-specific expert models, enabling efficient prediction of performance under various mixing strategies without repeated training. This method pioneers the use of model merging as a proxy mechanism for data mixture optimization, effectively decoupling performance evaluation from actual model training. Evaluated across 14 benchmarks, the surrogate model exhibits strong rank correlation with the true performance of mixed models, substantially improving search efficiency and scalability. The approach offers a low-cost, high-efficiency solution for optimizing multimodal fine-tuning pipelines.

Combinatorial SearchData Mixture OptimizationMixture Weights

Hot Scholars

ZL

Zerui Li

Adelaide Univeristy
RoboticsComputer VisionEmbodied AI
XS

Xiangyu Shi

Head of Algorithm Department, Metaradio
Applications of Deep LearningCompute LinguistCompute BiologyWireless Commnucation
CM

Chris McCool

Senior Principal Research Scientist, CSIRO
Computer VisionFine-Grained ClassificationAgricultural RoboticsPattern Recognition
YZ

Yue Zhao

Assistant Professor of Computer Science, University of Southern California
Anomaly DetectionOut-of-Distribution DetectionTrustworthy AIAI for Science
SH

Shuangping Huang

Professor, Electronic and Information Engineering, South China University of Technology
Computer VisionAIGCLLMEmbodied AI