Score
Design, build, or evaluate systems that apply general-purpose multimodal large language models (MLLMs) to process, interpret, or generate combined text and visual outputs. This work includes integrating off-the-shelf MLLMs into task pipelines, leveraging their cross-modal and transfer capabilities without task-specific fine-tuning, and measuring model performance, robustness, and transfer.
A systematic integration and critical reflection on evaluation frameworks for multimodal large language models (MLLMs) remains absent. Method: This paper presents the first meta-review of the MLLM field, conducting bibliometric analysis, topic modeling, and cross-review comparative analysis on 87 survey papers—performing structured metadata extraction and influence tracing across dimensions including benchmarking, methodology, application scenarios, ethical/safety considerations, and efficiency. Contribution/Results: It enables meta-level categorization and bias identification, exposing structural blind spots in existing surveys regarding coverage breadth, evaluation depth, and evolutionary tracking. The study distills seven core evaluation challenges and, for the first time, identifies three chronically underemphasized assessment dimensions: cross-modal causal reasoning, long-term temporal consistency, and multimodal grounding fidelity. These findings have been adopted as domain benchmarks by 12 subsequent works.
This paper addresses the challenges in developing general-purpose multimodal large language models (MLLMs) capable of cross-modal coordination across six generative modalities: text, image, music, video, human motion, and 3D objects. To this end, it proposes a novel unified architecture integrating Transformer-based and diffusion-based paradigms, augmented with self-supervised learning (SSL), mixture-of-experts (MoE), reinforcement learning from human feedback (RLHF), and chain-of-thought (CoT) reasoning. The work introduces the first taxonomy covering all six modalities and identifies shared enabling mechanisms for cross-modal transfer. It further argues that structured reasoning and modular decoupling are critical to improving interpretability and generalization. The resulting comprehensive MLLM technology landscape clarifies common bottlenecks and transferable methodologies, providing both theoretical foundations and practical guidelines for building universal, adaptive, and interpretable multimodal systems.
This paper addresses key challenges in multimodal large language models (MLLMs)—namely limited scalability, insufficient robustness, and difficulties in cross-modal alignment—within vision-language tasks. To tackle these, it proposes a comprehensive technical taxonomy covering the full MLLM stack and introduces a tripartite evaluation framework centered on scalability, robustness, and ethical alignment. Methodologically, the work conducts a systematic comparative analysis of core techniques, including Transformer-based fusion architectures, visual encoders (e.g., ViT, CLIP), cross-modal alignment learning, instruction tuning, and multi-stage training. It further synthesizes structured best practices for modeling across text, image, video, and audio modalities. The study delivers 12 representative application cases, empirically validating its framework and guidelines. Collectively, this work advances trustworthy deployment and cross-domain adoption of MLLMs through rigorous architectural analysis, principled evaluation, and actionable implementation insights.
Current multimodal large language model evaluation benchmarks exhibit significant blind spots in assessing cross-modal integration capabilities, particularly lacking coverage of critical cognitive dimensions such as spatiotemporal consistency, physical world understanding, multimodal coherence, and selective attention. This work addresses this gap by systematically reviewing existing evaluation methodologies and constructing a multimodal cognitive capability taxonomy, thereby explicitly identifying the core cognitive capacities missing from prevailing benchmarks. The study not only reveals substantial deficiencies in current evaluation frameworks but also provides a theoretical foundation and clear direction for developing more comprehensive and cognitively plausible multimodal assessment standards.
Current multimodal large language models (MLLMs) lack comprehensive, reproducible evaluation benchmarks, hindering accurate characterization of their multimodal understanding capabilities. To address this, we introduce MME—the first holistic benchmark for MLLM evaluation—comprising 14 fine-grained subtasks (e.g., OCR, visual reasoning, commonsense reasoning) that systematically assess both perceptual and cognitive abilities. All instruction-answer pairs are manually crafted to prevent data leakage and training contamination; a minimal, generic instruction template is employed to decouple intrinsic model capability from prompt engineering effects. MME provides a standardized evaluation protocol, an open-source dataset, and a public leaderboard. Applying MME to uniformly evaluate 30 state-of-the-art MLLMs reveals critical bottlenecks in cross-modal alignment and fine-grained perception, offering concrete insights for future research and development.
Current large language models (LLMs) rely solely on textual instructions for tool invocation, limiting their ability to accurately infer users’ underlying intentions—particularly under modality ambiguity and when multiple functionally equivalent tools are available. To address this, we propose MM-ToolLLM, the first multimodal LLM explicitly designed for tool calling, integrating vision (ViT) and audio (Whisper) encoders into a tool agent to enable cross-modal intent understanding and robust tool matching. We introduce the first multimodal tool-instruction dataset featuring multiple candidate tools per query, and adopt a joint training strategy combining instruction tuning with multimodal alignment. Experiments demonstrate significant improvements in tool recommendation accuracy, effective resolution of ambiguous queries, and support for multi-solution recommendations among functionally equivalent tools. The code and dataset are publicly released.
Multimodal large language models (MLLMs) suffer from pervasive hallucinations stemming from visual–linguistic semantic misalignment, severely undermining their reliability and real-world applicability. This work systematically investigates the root causes of such hallucinations and introduces, for the first time, a fine-grained taxonomy. We integrate mainstream benchmarks—including POPE and MME—with quantitative metrics to conduct cross-modal alignment diagnostics and empirical evaluation. Further, we synthesize and categorize mitigation strategies across three dimensions: prompt engineering, parameter-efficient fine-tuning, and decoding control, constructing a comprehensive methodological map. Our key contributions include: (1) a unified analytical framework for MLLM hallucination; (2) an open-source resource repository, Awesome-MLLM-Hallucination; and (3) a clear articulation of open challenges and future research directions—thereby providing both theoretical foundations and practical guidelines for enhancing MLLM robustness.
Multimodal Large Language Models (MLLMs) combine the natural language understanding and generation capabilities of LLMs with perception skills in modalities such as image and audio, representing a key advancement in contemporary AI. This chapter presents the main fundamentals of MLLMs and emblematic models. Practical techniques for preprocessing, prompt engineering, and building multimodal pipelines with LangChain and LangGraph are also explored. For further practical study, supplementary material is publicly available online: https://github.com/neemiasbsilva/MLLMs-Teoria-e-Pratica. Finally, the chapter discusses the challenges and highlights promising trends.
Multimodal large language models (MLLMs) face critical limitations in low-quality multimodal data, poor generalization on complex tasks, and inadequate evaluation frameworks—hindering their reliable deployment. This paper presents the first systematic survey of tool-augmented MLLMs, organized along four dimensions: data construction, task enhancement, evaluation methodology, and future challenges. We propose a unified collaborative framework integrating multimodal encoders, large language models, and external tools—including APIs, domain-specific expert models, and knowledge bases—to significantly improve cross-modal understanding and reasoning. Furthermore, we develop an open-source toolkit repository to empirically validate diverse tool-augmentation strategies. Our work provides both theoretical foundations and practical guidelines for high-fidelity multimodal data generation, robust complex-scenario problem solving, and trustworthy evaluation of MLLMs—advancing multimodal AI toward greater reliability, interpretability, and scalability.
This study investigates whether document information extraction in the era of multimodal large language models (MLLMs) still necessitates traditional OCR preprocessing. Through large-scale benchmarking, the authors evaluate the end-to-end performance of off-the-shelf MLLMs on real-world business documents and propose a direct image input pipeline that bypasses OCR entirely. They develop an LLM-based automated hierarchical error analysis framework to systematically diagnose failure modes and introduce structured prompting, in-context learning, and output schema constraints to significantly enhance extraction accuracy. Experimental results demonstrate that the OCR-free approach achieves comparable accuracy to OCR-augmented pipelines across most scenarios, confirming the feasibility of streamlining the extraction workflow and offering practical guidance for real-world deployment.
This paper addresses the critical challenge of enabling autonomous capability enhancement in multimodal large language models (MLLMs) under low-human-effort constraints. We present the first systematic survey of self-improvement mechanisms in MLLMs, proposing a three-dimensional framework encompassing data augmentation (generation, feedback integration, and filtering), data organization (curriculum learning, memory mechanisms, and reinforcement learning), and model optimization. Innovatively, we establish a hierarchical “Data–Organization–Optimization” taxonomy to unify mainstream methodologies, evaluation paradigms, and application scenarios. We explicitly identify key bottlenecks—including cross-modal alignment bias, heavy dependence on feedback quality, and limited generalizability—as open challenges. Our work delivers the first structured research blueprint for MLLM autonomous evolution, advancing the development of cost-efficient and sustainable multimodal agents.
This work addresses the lack of systematic evaluation of multimodal large language models (MLLMs) in real-world assistive AI scenarios by introducing NetraLink, a first-person vision system built upon head-mounted GoPro devices. The authors collect a benchmark dataset encompassing representative tasks such as everyday visual understanding, scene text question answering, and multilingual reading. For the first time, they conduct an end-to-end evaluation of mainstream MLLMs within a genuine assistive setting, revealing critical performance boundaries and limitations in object recognition, textual comprehension, and cross-lingual interaction. Furthermore, the study proposes a multimodal evaluation framework tailored for assistive technologies, offering diagnostic insights and actionable directions for deploying MLLMs effectively in real-world applications.