mllm evaluation

Designs and implements evaluation systems that use multimodal large language models (mllms) to analyze and score system outputs; this includes crafting prompts and rubrics, building automated scoring and aggregation pipelines, and validating measures of instruction compliance, fidelity, checklist criteria, and scalability across large item sets.

mllmevaluation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.75
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Jun 23, 2023
CF
Chaoyou Fu
🏛️ Tencent Youtu Lab | Xiamen University

Current multimodal large language models (MLLMs) lack comprehensive, reproducible evaluation benchmarks, hindering accurate characterization of their multimodal understanding capabilities. To address this, we introduce MME—the first holistic benchmark for MLLM evaluation—comprising 14 fine-grained subtasks (e.g., OCR, visual reasoning, commonsense reasoning) that systematically assess both perceptual and cognitive abilities. All instruction-answer pairs are manually crafted to prevent data leakage and training contamination; a minimal, generic instruction template is employed to decouple intrinsic model capability from prompt engineering effects. MME provides a standardized evaluation protocol, an open-source dataset, and a public leaderboard. Applying MME to uniformly evaluate 30 state-of-the-art MLLMs reveals critical bottlenecks in cross-modal alignment and fine-grained perception, offering concrete insights for future research and development.

Creating a comprehensive benchmark for multimodal large language modelsEvaluating 30 advanced MLLMs to identify improvement areasMeasuring perception and cognition across 14 multimodal subtasks

Despite rapid advances in multimodal large language models (MLLMs), their clinical deployment remains hindered by domain-specific bottlenecks—including scarce annotated medical data, modality bias, and limited interpretability. Method: This work systematically reviews the evolution from large language models (LLMs) to MLLMs and empirically analyzes their integration of text, medical imaging, and audio modalities for clinical decision support, radiology/pathology interpretation, patient interaction, and biomedical research. Contribution/Results: We identify three critical research directions: (1) construction of medical-domain multimodal datasets, (2) novel modality alignment techniques, and (3) an ethics-aware governance framework. Empirical evaluation demonstrates that MLLMs improve diagnostic assistance accuracy and accelerate structured reporting generation; however, performance is constrained by data scarcity, cross-modal misalignment, and opaque reasoning. This study provides both theoretical foundations and actionable pathways toward trustworthy, clinically viable MLLMs.

Addresses implementation challenges including data limitations and ethical concernsExamines MLLM applications in clinical decision support and medical imagingExplores the evolution of text-based LLMs to multimodal systems in healthcare

This paper addresses key challenges in multimodal large language models (MLLMs)—namely limited scalability, insufficient robustness, and difficulties in cross-modal alignment—within vision-language tasks. To tackle these, it proposes a comprehensive technical taxonomy covering the full MLLM stack and introduces a tripartite evaluation framework centered on scalability, robustness, and ethical alignment. Methodologically, the work conducts a systematic comparative analysis of core techniques, including Transformer-based fusion architectures, visual encoders (e.g., ViT, CLIP), cross-modal alignment learning, instruction tuning, and multi-stage training. It further synthesizes structured best practices for modeling across text, image, video, and audio modalities. The study delivers 12 representative application cases, empirically validating its framework and guidelines. Collectively, this work advances trustworthy deployment and cross-domain adoption of MLLMs through rigorous architectural analysis, principled evaluation, and actionable implementation insights.

Addressing challenges in scalability, robustness, and cross-modal learningExploring integration of text, images, video, and audio for cross-modal understandingSurveying multimodal large language models' architectures and applications

MLLM-Tool: A Multimodal Large Language Model for Tool Agent Learning

Jan 19, 2024
CW
Chenyu Wang
🏛️ ShanghaiTech University | Meituan | UniDT

Current large language models (LLMs) rely solely on textual instructions for tool invocation, limiting their ability to accurately infer users’ underlying intentions—particularly under modality ambiguity and when multiple functionally equivalent tools are available. To address this, we propose MM-ToolLLM, the first multimodal LLM explicitly designed for tool calling, integrating vision (ViT) and audio (Whisper) encoders into a tool agent to enable cross-modal intent understanding and robust tool matching. We introduce the first multimodal tool-instruction dataset featuring multiple candidate tools per query, and adopt a joint training strategy combining instruction tuning with multimodal alignment. Experiments demonstrate significant improvements in tool recommendation accuracy, effective resolution of ambiguous queries, and support for multi-solution recommendations among functionally equivalent tools. The code and dataset are publicly released.

Address ambiguity in user intentions with visual/auditory dataEnable multi-modal input for accurate tool selectionEnhance LLMs' tool perception beyond text queries

Existing evaluation benchmarks inadequately assess multimodal large language models’ (MLLMs) capability to follow complex, hierarchical instructions. Method: We introduce MIA-Bench—a rigorously constructed benchmark comprising 400 image-instruction pairs—and propose “instruction fidelity” as a novel, core evaluation dimension. We design structured instruction challenges and a fine-grained compliance assessment protocol. High-quality test samples are generated via human-crafted templates and pattern constraints; model adherence is improved through supervised fine-tuning and instruction-augmented data. Contribution/Results: Experiments reveal substantial performance gaps among state-of-the-art MLLMs on MIA-Bench. Targeted fine-tuning boosts average instruction compliance by 23.6% without degrading general vision-language capabilities. This work establishes a new paradigm for systematic evaluation and optimization of MLLM instruction-following behavior.

Benchmark with 400 challenging image-prompt pairsEnhance model instruction fidelity via trainingEvaluate multimodal LLMs instruction adherence

Latest Papers

What's happening recently
View more

Current multimodal large language model evaluation benchmarks exhibit significant blind spots in assessing cross-modal integration capabilities, particularly lacking coverage of critical cognitive dimensions such as spatiotemporal consistency, physical world understanding, multimodal coherence, and selective attention. This work addresses this gap by systematically reviewing existing evaluation methodologies and constructing a multimodal cognitive capability taxonomy, thereby explicitly identifying the core cognitive capacities missing from prevailing benchmarks. The study not only reveals substantial deficiencies in current evaluation frameworks but also provides a theoretical foundation and clear direction for developing more comprehensive and cognitively plausible multimodal assessment standards.

multimodal consistencymultimodal evaluationphysical world understanding

Multimodal Large Language Models (MLLMs) combine the natural language understanding and generation capabilities of LLMs with perception skills in modalities such as image and audio, representing a key advancement in contemporary AI. This chapter presents the main fundamentals of MLLMs and emblematic models. Practical techniques for preprocessing, prompt engineering, and building multimodal pipelines with LangChain and LangGraph are also explored. For further practical study, supplementary material is publicly available online: https://github.com/neemiasbsilva/MLLMs-Teoria-e-Pratica. Finally, the chapter discusses the challenges and highlights promising trends.

AI integrationMultimodal Large Language Modelsmultimodal perception

Existing security benchmarks lack academic domain specificity and fail to capture the dynamic phishing threats and human cognitive vulnerabilities faced by multimodal large language models (MLLMs) in scholarly contexts. Method: We propose AdapT-Bench—the first multilingual, contextualized, multimodal phishing detection benchmark explicitly grounded in academic knowledge—supporting dynamic attack simulation, cross-lingual threat modeling, and human cognitive vulnerability assessment. Our approach integrates multimodal reasoning, context-aware analysis, and controllable data generation to enable fine-grained identification of highly customized phishing content. Contribution/Results: Experiments demonstrate that AdapT-Bench significantly improves MLLMs’ phishing detection accuracy in academic settings, empirically validating the effectiveness and necessity of academic knowledge injection and joint multimodal-contextual modeling.

Addressing inadequacy of existing benchmarks for academic-specific phishing attacksDeveloping comprehensive framework to test multimodal phishing detection capabilitiesEvaluating MLLM security against dynamic phishing threats in academic environments

This study investigates whether document information extraction in the era of multimodal large language models (MLLMs) still necessitates traditional OCR preprocessing. Through large-scale benchmarking, the authors evaluate the end-to-end performance of off-the-shelf MLLMs on real-world business documents and propose a direct image input pipeline that bypasses OCR entirely. They develop an LLM-based automated hierarchical error analysis framework to systematically diagnose failure modes and introduce structured prompting, in-context learning, and output schema constraints to significantly enhance extraction accuracy. Experimental results demonstrate that the OCR-free approach achieves comparable accuracy to OCR-augmented pipelines across most scenarios, confirming the feasibility of streamlining the extraction workflow and offering practical guidance for real-world deployment.

Business DocumentsDocument Information ExtractionMultimodal Large Language Models

Seeing the Big Picture: Evaluating Multimodal LLMs' Ability to Interpret and Grade Handwritten Student Work

Oct 06, 2025
OH
Owen Henkel
🏛️ University of Oxford | Legible Labs | Coherence Fund | XQ Institute

This study systematically evaluates multimodal large language models (MLLMs) for automated parsing and grading of elementary school handwritten mathematics assignments, focusing on arithmetic answer recognition and open-ended mathematical diagram evaluation. To mitigate MLLMs’ overreliance on image quality, we propose a novel “description-augmented” paradigm that incorporates human-generated semantic descriptions as auxiliary textual input alongside visual inputs. Experimental results show that, on objective arithmetic problems, MLLMs achieve 95% accuracy (Cohen’s κ = 0.90) when processing raw images—approaching human-level performance. In contrast, for diagram evaluation, inter-rater agreement drops to κ = 0.20 with image-only input but improves significantly to κ = 0.47 with description augmentation—matching the consistency observed among human graders. This work is the first to empirically reveal a substantial capability gap between MLLMs’ performance on objective versus open-ended handwritten mathematics tasks, and demonstrates that text-based augmentation substantially enhances both reliability and interpretability in open-ended educational assessment.

Assessing MLLM performance on mathematical illustrations and arithmeticEvaluating MLLMs' ability to grade handwritten student workInvestigating gaps in visual interpretation versus pedagogical judgment

Hot Scholars

YZ

Yuanxing Zhang

Kuaishou Technology
Recommender SystemLarge Language ModelVideo Understanding
YT

Yuqi Tang

Duke University
Medical ImagingComputer VisionImage Quality
YS

Yiyang Su

Michigan State University
Computer Vision
XL

Xiaoming Liu

Anil K. and Nandita Jain Endowed Professor; MSU Foundation Professor, CSE, Michigan State University
Computer VisionPattern RecognitionMachine LearningBiometrics
YD

Yuhao Dong

Tsinghua University, Nanyang Technological University
Multi-modal LearningComputer Vision