general-purpose mllms

Design, build, or evaluate systems that apply general-purpose multimodal large language models (MLLMs) to process, interpret, or generate combined text and visual outputs. This work includes integrating off-the-shelf MLLMs into task pipelines, leveraging their cross-modal and transfer capabilities without task-specific fine-tuning, and measuring model performance, robustness, and transfer.

general-purposemllms

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-1.45
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Multimodal Large Language Models: A Survey

May 29, 2025
LH
Longzhen Han
🏛️ University of Brighton

This paper addresses the challenges in developing general-purpose multimodal large language models (MLLMs) capable of cross-modal coordination across six generative modalities: text, image, music, video, human motion, and 3D objects. To this end, it proposes a novel unified architecture integrating Transformer-based and diffusion-based paradigms, augmented with self-supervised learning (SSL), mixture-of-experts (MoE), reinforcement learning from human feedback (RLHF), and chain-of-thought (CoT) reasoning. The work introduces the first taxonomy covering all six modalities and identifies shared enabling mechanisms for cross-modal transfer. It further argues that structured reasoning and modular decoupling are critical to improving interpretability and generalization. The resulting comprehensive MLLM technology landscape clarifies common bottlenecks and transferable methodologies, providing both theoretical foundations and practical guidelines for building universal, adaptive, and interpretable multimodal systems.

Examines techniques enabling cross-modal capabilitiesIdentifies challenges in evaluation and modularitySurvey categorizes generative modalities in MLLMs

Must-Read Papers

Most classic and influential ideas
View more

This paper addresses key challenges in multimodal large language models (MLLMs)—namely limited scalability, insufficient robustness, and difficulties in cross-modal alignment—within vision-language tasks. To tackle these, it proposes a comprehensive technical taxonomy covering the full MLLM stack and introduces a tripartite evaluation framework centered on scalability, robustness, and ethical alignment. Methodologically, the work conducts a systematic comparative analysis of core techniques, including Transformer-based fusion architectures, visual encoders (e.g., ViT, CLIP), cross-modal alignment learning, instruction tuning, and multi-stage training. It further synthesizes structured best practices for modeling across text, image, video, and audio modalities. The study delivers 12 representative application cases, empirically validating its framework and guidelines. Collectively, this work advances trustworthy deployment and cross-domain adoption of MLLMs through rigorous architectural analysis, principled evaluation, and actionable implementation insights.

Addressing challenges in scalability, robustness, and cross-modal learningExploring integration of text, images, video, and audio for cross-modal understandingSurveying multimodal large language models' architectures and applications

Current multimodal large language model evaluation benchmarks exhibit significant blind spots in assessing cross-modal integration capabilities, particularly lacking coverage of critical cognitive dimensions such as spatiotemporal consistency, physical world understanding, multimodal coherence, and selective attention. This work addresses this gap by systematically reviewing existing evaluation methodologies and constructing a multimodal cognitive capability taxonomy, thereby explicitly identifying the core cognitive capacities missing from prevailing benchmarks. The study not only reveals substantial deficiencies in current evaluation frameworks but also provides a theoretical foundation and clear direction for developing more comprehensive and cognitively plausible multimodal assessment standards.

multimodal consistencymultimodal evaluationphysical world understanding

MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Jun 23, 2023
CF
Chaoyou Fu
🏛️ Tencent Youtu Lab | Xiamen University

Current multimodal large language models (MLLMs) lack comprehensive, reproducible evaluation benchmarks, hindering accurate characterization of their multimodal understanding capabilities. To address this, we introduce MME—the first holistic benchmark for MLLM evaluation—comprising 14 fine-grained subtasks (e.g., OCR, visual reasoning, commonsense reasoning) that systematically assess both perceptual and cognitive abilities. All instruction-answer pairs are manually crafted to prevent data leakage and training contamination; a minimal, generic instruction template is employed to decouple intrinsic model capability from prompt engineering effects. MME provides a standardized evaluation protocol, an open-source dataset, and a public leaderboard. Applying MME to uniformly evaluate 30 state-of-the-art MLLMs reveals critical bottlenecks in cross-modal alignment and fine-grained perception, offering concrete insights for future research and development.

Creating a comprehensive benchmark for multimodal large language modelsEvaluating 30 advanced MLLMs to identify improvement areasMeasuring perception and cognition across 14 multimodal subtasks

MLLM-Tool: A Multimodal Large Language Model for Tool Agent Learning

Jan 19, 2024
CW
Chenyu Wang
🏛️ ShanghaiTech University | Meituan | UniDT

Current large language models (LLMs) rely solely on textual instructions for tool invocation, limiting their ability to accurately infer users’ underlying intentions—particularly under modality ambiguity and when multiple functionally equivalent tools are available. To address this, we propose MM-ToolLLM, the first multimodal LLM explicitly designed for tool calling, integrating vision (ViT) and audio (Whisper) encoders into a tool agent to enable cross-modal intent understanding and robust tool matching. We introduce the first multimodal tool-instruction dataset featuring multiple candidate tools per query, and adopt a joint training strategy combining instruction tuning with multimodal alignment. Experiments demonstrate significant improvements in tool recommendation accuracy, effective resolution of ambiguous queries, and support for multi-solution recommendations among functionally equivalent tools. The code and dataset are publicly released.

Address ambiguity in user intentions with visual/auditory dataEnable multi-modal input for accurate tool selectionEnhance LLMs' tool perception beyond text queries

Hallucination of Multimodal Large Language Models: A Survey

Apr 29, 2024
ZB
Zechen Bai
🏛️ National University of Singapore | Amazon Prime Video | AWS

Multimodal large language models (MLLMs) suffer from pervasive hallucinations stemming from visual–linguistic semantic misalignment, severely undermining their reliability and real-world applicability. This work systematically investigates the root causes of such hallucinations and introduces, for the first time, a fine-grained taxonomy. We integrate mainstream benchmarks—including POPE and MME—with quantitative metrics to conduct cross-modal alignment diagnostics and empirical evaluation. Further, we synthesize and categorize mitigation strategies across three dimensions: prompt engineering, parameter-efficient fine-tuning, and decoding control, constructing a comprehensive methodological map. Our key contributions include: (1) a unified analytical framework for MLLM hallucination; (2) an open-source resource repository, Awesome-MLLM-Hallucination; and (3) a clear articulation of open challenges and future research directions—thereby providing both theoretical foundations and practical guidelines for enhancing MLLM robustness.

Analyzing hallucination in multimodal large language models (MLLMs).Detecting and mitigating inconsistent outputs with visual content.Reviewing causes, benchmarks, and solutions for MLLM hallucinations.

Latest Papers

What's happening recently
View more

Multimodal Large Language Models (MLLMs) combine the natural language understanding and generation capabilities of LLMs with perception skills in modalities such as image and audio, representing a key advancement in contemporary AI. This chapter presents the main fundamentals of MLLMs and emblematic models. Practical techniques for preprocessing, prompt engineering, and building multimodal pipelines with LangChain and LangGraph are also explored. For further practical study, supplementary material is publicly available online: https://github.com/neemiasbsilva/MLLMs-Teoria-e-Pratica. Finally, the chapter discusses the challenges and highlights promising trends.

AI integrationMultimodal Large Language Modelsmultimodal perception

Empowering Multimodal LLMs with External Tools: A Comprehensive Survey

Aug 14, 2025
WA
Wenbin An
🏛️ Xi'an Jiaotong University | Nanyang Technological University | Lenovo Research | Lenovo | School of Computer Science and Technology | College of Computing and Data Science

Multimodal large language models (MLLMs) face critical limitations in low-quality multimodal data, poor generalization on complex tasks, and inadequate evaluation frameworks—hindering their reliable deployment. This paper presents the first systematic survey of tool-augmented MLLMs, organized along four dimensions: data construction, task enhancement, evaluation methodology, and future challenges. We propose a unified collaborative framework integrating multimodal encoders, large language models, and external tools—including APIs, domain-specific expert models, and knowledge bases—to significantly improve cross-modal understanding and reasoning. Furthermore, we develop an open-source toolkit repository to empirically validate diverse tool-augmentation strategies. Our work provides both theoretical foundations and practical guidelines for high-fidelity multimodal data generation, robust complex-scenario problem solving, and trustworthy evaluation of MLLMs—advancing multimodal AI toward greater reliability, interpretability, and scalability.

Addressing evaluation challenges in MLLM reliability and applicabilityEnhancing MLLMs with external tools for better performanceImproving multimodal data quality and annotation processes

This study investigates whether document information extraction in the era of multimodal large language models (MLLMs) still necessitates traditional OCR preprocessing. Through large-scale benchmarking, the authors evaluate the end-to-end performance of off-the-shelf MLLMs on real-world business documents and propose a direct image input pipeline that bypasses OCR entirely. They develop an LLM-based automated hierarchical error analysis framework to systematically diagnose failure modes and introduce structured prompting, in-context learning, and output schema constraints to significantly enhance extraction accuracy. Experimental results demonstrate that the OCR-free approach achieves comparable accuracy to OCR-augmented pipelines across most scenarios, confirming the feasibility of streamlining the extraction workflow and offering practical guidance for real-world deployment.

Business DocumentsDocument Information ExtractionMultimodal Large Language Models

Self-Improvement in Multimodal Large Language Models: A Survey

Oct 02, 2025
SD
Shijian Deng
🏛️ The University of Texas at Dallas | University of Toronto | University of Notre Dame | Mohamed bin Zayed University of Artificial Intelligence

This paper addresses the critical challenge of enabling autonomous capability enhancement in multimodal large language models (MLLMs) under low-human-effort constraints. We present the first systematic survey of self-improvement mechanisms in MLLMs, proposing a three-dimensional framework encompassing data augmentation (generation, feedback integration, and filtering), data organization (curriculum learning, memory mechanisms, and reinforcement learning), and model optimization. Innovatively, we establish a hierarchical “Data–Organization–Optimization” taxonomy to unify mainstream methodologies, evaluation paradigms, and application scenarios. We explicitly identify key bottlenecks—including cross-modal alignment bias, heavy dependence on feedback quality, and limited generalizability—as open challenges. Our work delivers the first structured research blueprint for MLLM autonomous evolution, advancing the development of cost-efficient and sustainable multimodal agents.

Identifying challenges and future directions for MLLM developmentStructuring literature on data collection and model optimizationSurveying self-improvement methods for Multimodal Large Language Models

This work addresses the lack of systematic evaluation of multimodal large language models (MLLMs) in real-world assistive AI scenarios by introducing NetraLink, a first-person vision system built upon head-mounted GoPro devices. The authors collect a benchmark dataset encompassing representative tasks such as everyday visual understanding, scene text question answering, and multilingual reading. For the first time, they conduct an end-to-end evaluation of mainstream MLLMs within a genuine assistive setting, revealing critical performance boundaries and limitations in object recognition, textual comprehension, and cross-lingual interaction. Furthermore, the study proposes a multimodal evaluation framework tailored for assistive technologies, offering diagnostic insights and actionable directions for deploying MLLMs effectively in real-world applications.

Assistive AIegocentric visionMultimodal Large Language Models

Hot Scholars

YL

Yangfu Li

East China Normal University
Deep learning
IG

Igor Gilitschenski

Assistant Professor, University of Toronto
RoboticsMachine LearningComputer Vision
YJ

Yu-Jhe Li

Research Scientist, Adobe
Computer VisionMachine LearningMachine PerceptionGenerative AI
MM

Michele Magno

ETH Zurich
Wireless sensor networksSmart Sensors and Internet of ThingsWake up RadioPower management