Score
Designs, builds, and evaluates large language model systems that jointly process and generate multiple data modalities (e.g., text plus images, audio, or video), including multimodal encoders/decoders, cross‑modal fusion and alignment modules, and inference pipelines. Creates and curates multimodal training and evaluation datasets, implements fine‑tuning and scaling workflows, and analyzes model performance, robustness, and cross‑modal behavior for integrated multimodal applications.
This paper addresses the challenges in developing general-purpose multimodal large language models (MLLMs) capable of cross-modal coordination across six generative modalities: text, image, music, video, human motion, and 3D objects. To this end, it proposes a novel unified architecture integrating Transformer-based and diffusion-based paradigms, augmented with self-supervised learning (SSL), mixture-of-experts (MoE), reinforcement learning from human feedback (RLHF), and chain-of-thought (CoT) reasoning. The work introduces the first taxonomy covering all six modalities and identifies shared enabling mechanisms for cross-modal transfer. It further argues that structured reasoning and modular decoupling are critical to improving interpretability and generalization. The resulting comprehensive MLLM technology landscape clarifies common bottlenecks and transferable methodologies, providing both theoretical foundations and practical guidelines for building universal, adaptive, and interpretable multimodal systems.
A standardized, integrated benchmark for training and evaluating multimodal large language models (MLLMs) is lacking, hindering systematic research on cross-modal model merging. Method: We introduce the first dedicated MLLM model merging benchmark, covering diverse tasks—including visual question answering, geometric reasoning, chart understanding, OCR, and localization—and supporting both LoRA and full-parameter fine-tuned models. We systematically investigate vision–language, audio–language, and video–language fusion pathways. Furthermore, we propose a novel fusion method integrating task vector denoising, interactive loss optimization, LoRA adapter fusion, and multimodal alignment. Results: Our approach achieves an average performance gain of 2.48% across benchmarks, surpassing unimodal expert models without additional training data. This constitutes the first empirical demonstration that multimodal complementarity can yield stronger, general-purpose Omni-language models.
This paper addresses key challenges in multimodal large language models (MLLMs)—namely limited scalability, insufficient robustness, and difficulties in cross-modal alignment—within vision-language tasks. To tackle these, it proposes a comprehensive technical taxonomy covering the full MLLM stack and introduces a tripartite evaluation framework centered on scalability, robustness, and ethical alignment. Methodologically, the work conducts a systematic comparative analysis of core techniques, including Transformer-based fusion architectures, visual encoders (e.g., ViT, CLIP), cross-modal alignment learning, instruction tuning, and multi-stage training. It further synthesizes structured best practices for modeling across text, image, video, and audio modalities. The study delivers 12 representative application cases, empirically validating its framework and guidelines. Collectively, this work advances trustworthy deployment and cross-domain adoption of MLLMs through rigorous architectural analysis, principled evaluation, and actionable implementation insights.
Multimodal large language models (LMMs) suffer from unstable inference performance, high resource consumption, and severe cross-modal request interference during production deployment. Method: We systematically analyze multi-stage inference behavior and resource contention patterns across six open-source LMMs, comparing decoder-only and cross-attention architectures. We propose a decoupled serving architecture enabling per-stage independent resource allocation and elastic scaling, and introduce stage co-location optimization—jointly maximizing throughput and resource utilization under latency constraints. Our approach integrates multimodal request trajectory modeling, heterogeneous resource scheduling, and compute-memory co-optimization. Contribution/Results: Evaluation shows up to 2.3× higher throughput, 41% reduction in tail latency, and 37% improvement in resource utilization compared to state-of-the-art baselines.
Existing research lacks a systematic analysis of modality–language backbone integration mechanisms in multimodal large language models (MLLMs). Method: We systematically survey 125 MLLMs published between 2021 and 2025, proposing the first large-language-model-centric three-dimensional taxonomy—architectural integration, representation learning, and training paradigms—to unify cross-modal alignment pathways. Leveraging bibliometric analysis, architectural decoupling, and fine-grained modeling of embeddings and loss functions, we characterize evolutionary patterns in fusion granularity, joint representation design, and objective function development. Contribution/Results: This work fills a critical theoretical gap in structured analysis of MLLM fusion mechanisms, identifies shared bottlenecks—including misaligned modality-specific feature hierarchies and suboptimal cross-modal optimization objectives—and delivers a reusable theoretical framework and practical guidelines for developing robust, scalable multimodal foundation models.
Existing open-source multimodal large language models (MLLMs) exhibit significant deficiencies in joint visual-auditory-textual understanding and reasoning, achieving only ~50% instruction-following accuracy on trilingual multimodal tasks. Method: We introduce OmniBench—the first benchmark for trilingual multimodal collaborative reasoning—and formalize the omni-language model (OLM), a unified architecture capable of jointly processing visual, auditory, and textual (V-A-T) inputs. We construct OmniBench via expert human annotation across diverse trilingual multimodal tasks and curate OmniInstruct, a large-scale instruction-tuning dataset comprising 96K samples. Our methodology integrates cross-modal alignment modeling, trilingual multimodal instruction tuning, and a human-in-the-loop evaluation framework. Contribution/Results: Experiments reveal severe generalization limitations of current open-source OLMs on trilingual multimodal tasks; OmniInstruct substantially improves their reasoning performance. This work establishes a novel evaluation paradigm, provides high-quality resources, and outlines a technical pathway for advancing trilingual multimodal foundation models.
Existing model-serving frameworks struggle to efficiently support increasingly complex composite multimodal models. To address this challenge, this work proposes M*, a novel system that introduces a modular Walk Graph abstraction to uniformly represent composite AI models as dataflow graphs. This abstraction enables flexible composition of arbitrary model components, cluster deployment, and model-agnostic distributed runtime optimizations. By leveraging a graph traversal mechanism, M* efficiently handles cross-modal, multitask requests, achieving significant performance gains: it reduces end-to-end latency by 20% over vLLM-Omni in text-to-image generation, improves real-time factor by 2.9× and throughput by 2.7× in text-to-speech tasks, and accelerates robot planning workloads by up to 12.5×.
This work addresses the limitations of existing approaches that predominantly rely on non-native late-fusion architectures and lack a systematic definition or unified framework for native multimodal modeling. The paper proposes a formal roadmap for Native Multimodal Modeling (NMM), introducing, for the first time, a rigorous formulation of “architectural nativeness.” Leveraging an input–output duality perspective, it categorizes models into three types: Multi-to-Text, Multi-to-Target, and Multi-to-Multi. A full-stack, industrial-grade NMM framework is developed, encompassing data governance, early/mid-stage fusion, end-to-end training, and deployment, all realized through a unified Transformer architecture that enables symbiotic cross-modal understanding and generation. This paradigm demonstrates superior performance and strong scalability across multiple tasks, offering a clear pathway toward truly native multimodal models.
This study addresses the unclear trade-offs among diverse capabilities—such as general visual understanding, structured reasoning, and fine-grained OCR—in multimodal instruction tuning under mixed data regimes, particularly the lack of systematic investigation into how data organization influences these trade-offs. Treating data scheduling as a first-order design variable while holding model architecture and optimization settings fixed, the work compares four strategies: direct mixing, curriculum learning, balanced sampling, and reverse curriculum. Results demonstrate that curriculum-based training—sequencing tasks from general comprehension to specialized skills—achieves superior overall performance and structured reasoning while accelerating convergence. Balanced sampling improves OCR accuracy at the cost of capability imbalance, whereas reverse curriculum degrades performance and induces optimization instability. The findings highlight the critical role of training sequence in shaping the capability distribution of multimodal models.
Multimodal Large Language Models (MLLMs) combine the natural language understanding and generation capabilities of LLMs with perception skills in modalities such as image and audio, representing a key advancement in contemporary AI. This chapter presents the main fundamentals of MLLMs and emblematic models. Practical techniques for preprocessing, prompt engineering, and building multimodal pipelines with LangChain and LangGraph are also explored. For further practical study, supplementary material is publicly available online: https://github.com/neemiasbsilva/MLLMs-Teoria-e-Pratica. Finally, the chapter discusses the challenges and highlights promising trends.
This work addresses the high cost and combinatorial explosion associated with optimizing data mixing ratios in cross-domain supervised fine-tuning of multimodal large language models. The authors propose an innovative approach that constructs a surrogate model by linearly merging parameters from domain-specific expert models, enabling efficient prediction of performance under various mixing strategies without repeated training. This method pioneers the use of model merging as a proxy mechanism for data mixture optimization, effectively decoupling performance evaluation from actual model training. Evaluated across 14 benchmarks, the surrogate model exhibits strong rank correlation with the true performance of mixed models, substantially improving search efficiency and scalability. The approach offers a low-cost, high-efficiency solution for optimizing multimodal fine-tuning pipelines.