Score
Designs and implements evaluation systems that use multimodal large language models (mllms) to analyze and score system outputs; this includes crafting prompts and rubrics, building automated scoring and aggregation pipelines, and validating measures of instruction compliance, fidelity, checklist criteria, and scalability across large item sets.
Current multimodal large language models (MLLMs) lack comprehensive, reproducible evaluation benchmarks, hindering accurate characterization of their multimodal understanding capabilities. To address this, we introduce MME—the first holistic benchmark for MLLM evaluation—comprising 14 fine-grained subtasks (e.g., OCR, visual reasoning, commonsense reasoning) that systematically assess both perceptual and cognitive abilities. All instruction-answer pairs are manually crafted to prevent data leakage and training contamination; a minimal, generic instruction template is employed to decouple intrinsic model capability from prompt engineering effects. MME provides a standardized evaluation protocol, an open-source dataset, and a public leaderboard. Applying MME to uniformly evaluate 30 state-of-the-art MLLMs reveals critical bottlenecks in cross-modal alignment and fine-grained perception, offering concrete insights for future research and development.
Despite rapid advances in multimodal large language models (MLLMs), their clinical deployment remains hindered by domain-specific bottlenecks—including scarce annotated medical data, modality bias, and limited interpretability. Method: This work systematically reviews the evolution from large language models (LLMs) to MLLMs and empirically analyzes their integration of text, medical imaging, and audio modalities for clinical decision support, radiology/pathology interpretation, patient interaction, and biomedical research. Contribution/Results: We identify three critical research directions: (1) construction of medical-domain multimodal datasets, (2) novel modality alignment techniques, and (3) an ethics-aware governance framework. Empirical evaluation demonstrates that MLLMs improve diagnostic assistance accuracy and accelerate structured reporting generation; however, performance is constrained by data scarcity, cross-modal misalignment, and opaque reasoning. This study provides both theoretical foundations and actionable pathways toward trustworthy, clinically viable MLLMs.
This paper addresses key challenges in multimodal large language models (MLLMs)—namely limited scalability, insufficient robustness, and difficulties in cross-modal alignment—within vision-language tasks. To tackle these, it proposes a comprehensive technical taxonomy covering the full MLLM stack and introduces a tripartite evaluation framework centered on scalability, robustness, and ethical alignment. Methodologically, the work conducts a systematic comparative analysis of core techniques, including Transformer-based fusion architectures, visual encoders (e.g., ViT, CLIP), cross-modal alignment learning, instruction tuning, and multi-stage training. It further synthesizes structured best practices for modeling across text, image, video, and audio modalities. The study delivers 12 representative application cases, empirically validating its framework and guidelines. Collectively, this work advances trustworthy deployment and cross-domain adoption of MLLMs through rigorous architectural analysis, principled evaluation, and actionable implementation insights.
Current large language models (LLMs) rely solely on textual instructions for tool invocation, limiting their ability to accurately infer users’ underlying intentions—particularly under modality ambiguity and when multiple functionally equivalent tools are available. To address this, we propose MM-ToolLLM, the first multimodal LLM explicitly designed for tool calling, integrating vision (ViT) and audio (Whisper) encoders into a tool agent to enable cross-modal intent understanding and robust tool matching. We introduce the first multimodal tool-instruction dataset featuring multiple candidate tools per query, and adopt a joint training strategy combining instruction tuning with multimodal alignment. Experiments demonstrate significant improvements in tool recommendation accuracy, effective resolution of ambiguous queries, and support for multi-solution recommendations among functionally equivalent tools. The code and dataset are publicly released.
Existing evaluation benchmarks inadequately assess multimodal large language models’ (MLLMs) capability to follow complex, hierarchical instructions. Method: We introduce MIA-Bench—a rigorously constructed benchmark comprising 400 image-instruction pairs—and propose “instruction fidelity” as a novel, core evaluation dimension. We design structured instruction challenges and a fine-grained compliance assessment protocol. High-quality test samples are generated via human-crafted templates and pattern constraints; model adherence is improved through supervised fine-tuning and instruction-augmented data. Contribution/Results: Experiments reveal substantial performance gaps among state-of-the-art MLLMs on MIA-Bench. Targeted fine-tuning boosts average instruction compliance by 23.6% without degrading general vision-language capabilities. This work establishes a new paradigm for systematic evaluation and optimization of MLLM instruction-following behavior.
Current multimodal large language model evaluation benchmarks exhibit significant blind spots in assessing cross-modal integration capabilities, particularly lacking coverage of critical cognitive dimensions such as spatiotemporal consistency, physical world understanding, multimodal coherence, and selective attention. This work addresses this gap by systematically reviewing existing evaluation methodologies and constructing a multimodal cognitive capability taxonomy, thereby explicitly identifying the core cognitive capacities missing from prevailing benchmarks. The study not only reveals substantial deficiencies in current evaluation frameworks but also provides a theoretical foundation and clear direction for developing more comprehensive and cognitively plausible multimodal assessment standards.
Multimodal Large Language Models (MLLMs) combine the natural language understanding and generation capabilities of LLMs with perception skills in modalities such as image and audio, representing a key advancement in contemporary AI. This chapter presents the main fundamentals of MLLMs and emblematic models. Practical techniques for preprocessing, prompt engineering, and building multimodal pipelines with LangChain and LangGraph are also explored. For further practical study, supplementary material is publicly available online: https://github.com/neemiasbsilva/MLLMs-Teoria-e-Pratica. Finally, the chapter discusses the challenges and highlights promising trends.
Existing security benchmarks lack academic domain specificity and fail to capture the dynamic phishing threats and human cognitive vulnerabilities faced by multimodal large language models (MLLMs) in scholarly contexts. Method: We propose AdapT-Bench—the first multilingual, contextualized, multimodal phishing detection benchmark explicitly grounded in academic knowledge—supporting dynamic attack simulation, cross-lingual threat modeling, and human cognitive vulnerability assessment. Our approach integrates multimodal reasoning, context-aware analysis, and controllable data generation to enable fine-grained identification of highly customized phishing content. Contribution/Results: Experiments demonstrate that AdapT-Bench significantly improves MLLMs’ phishing detection accuracy in academic settings, empirically validating the effectiveness and necessity of academic knowledge injection and joint multimodal-contextual modeling.
This study investigates whether document information extraction in the era of multimodal large language models (MLLMs) still necessitates traditional OCR preprocessing. Through large-scale benchmarking, the authors evaluate the end-to-end performance of off-the-shelf MLLMs on real-world business documents and propose a direct image input pipeline that bypasses OCR entirely. They develop an LLM-based automated hierarchical error analysis framework to systematically diagnose failure modes and introduce structured prompting, in-context learning, and output schema constraints to significantly enhance extraction accuracy. Experimental results demonstrate that the OCR-free approach achieves comparable accuracy to OCR-augmented pipelines across most scenarios, confirming the feasibility of streamlining the extraction workflow and offering practical guidance for real-world deployment.
This study systematically evaluates multimodal large language models (MLLMs) for automated parsing and grading of elementary school handwritten mathematics assignments, focusing on arithmetic answer recognition and open-ended mathematical diagram evaluation. To mitigate MLLMs’ overreliance on image quality, we propose a novel “description-augmented” paradigm that incorporates human-generated semantic descriptions as auxiliary textual input alongside visual inputs. Experimental results show that, on objective arithmetic problems, MLLMs achieve 95% accuracy (Cohen’s κ = 0.90) when processing raw images—approaching human-level performance. In contrast, for diagram evaluation, inter-rater agreement drops to κ = 0.20 with image-only input but improves significantly to κ = 0.47 with description augmentation—matching the consistency observed among human graders. This work is the first to empirically reveal a substantial capability gap between MLLMs’ performance on objective versus open-ended handwritten mathematics tasks, and demonstrates that text-based augmentation substantially enhances both reliability and interpretability in open-ended educational assessment.