Score
Designs, builds, and evaluates models and pipelines that combine visual inputs and natural language (multimodal LLMs) to infer, generate, or align step-by-step procedures and action sequences from observations and instructions. This includes mapping visual observations to procedural steps, inferring high-level task contexts and progression, and producing structured procedural outputs or contextual signals for downstream modules.
This paper addresses key challenges in multimodal large language models (MLLMs)—namely limited scalability, insufficient robustness, and difficulties in cross-modal alignment—within vision-language tasks. To tackle these, it proposes a comprehensive technical taxonomy covering the full MLLM stack and introduces a tripartite evaluation framework centered on scalability, robustness, and ethical alignment. Methodologically, the work conducts a systematic comparative analysis of core techniques, including Transformer-based fusion architectures, visual encoders (e.g., ViT, CLIP), cross-modal alignment learning, instruction tuning, and multi-stage training. It further synthesizes structured best practices for modeling across text, image, video, and audio modalities. The study delivers 12 representative application cases, empirically validating its framework and guidelines. Collectively, this work advances trustworthy deployment and cross-domain adoption of MLLMs through rigorous architectural analysis, principled evaluation, and actionable implementation insights.
This work evaluates the capability of current open-source multimodal large language models (MLMs) to provide real-time assistance in technical procedural tasks, with a focus on their limitations in understanding assembly instructions and aligning textual manuals with corresponding video actions. To this end, we introduce the first fine-grained Manual-to-Action Dataset (M2AD), which precisely aligns furniture assembly videos with their respective instruction steps. Leveraging M2AD, we systematically assess MLMs on key competencies including procedural comprehension, step tracking, and cross-modal referencing between text and images. Our findings reveal that, despite some models demonstrating preliminary procedural reasoning abilities, architectural and hardware constraints hinder their effective processing of multiple image inputs and complex interleaved text–image reasoning, thereby exposing significant gaps in their applicability to real-world technical assistance scenarios.
Existing multimodal large language models (MLLMs) employ only unidirectional mapping of visual information into the language space, failing to effectively leverage visual knowledge to enhance holistic reasoning. This work introduces Vision-Augmented Large Language Models (VA-LLMs), breaking from this paradigm by enabling LLMs to actively store, share, and retrieve visual knowledge. Our contributions are threefold: (1) Modular Visual Memory (MVM), a structured, long-term storage mechanism for visual knowledge; (2) Soft Multimodal Mixture of Experts (MoME), a dynamic architecture that coordinates vision and language experts during autoregressive generation; and (3) a cross-modal knowledge injection and collaborative reasoning framework. Experiments demonstrate substantial improvements in physical commonsense understanding and spatial reasoning, achieving state-of-the-art performance across multiple multimodal benchmarks, including MMMU, ScienceQA, and POPE.
Multimodal Large Language Models (MLLMs) combine the natural language understanding and generation capabilities of LLMs with perception skills in modalities such as image and audio, representing a key advancement in contemporary AI. This chapter presents the main fundamentals of MLLMs and emblematic models. Practical techniques for preprocessing, prompt engineering, and building multimodal pipelines with LangChain and LangGraph are also explored. For further practical study, supplementary material is publicly available online: https://github.com/neemiasbsilva/MLLMs-Teoria-e-Pratica. Finally, the chapter discusses the challenges and highlights promising trends.
Existing video procedural planning methods rely on LLMs to generate fixed action sequences, suffering from poor generalization and limited adaptability to novel tasks or open-vocabulary actions. This work addresses video program planning for embodied intelligence—i.e., inferring executable action sequences from given start and goal frames. We propose an LLM-enhanced planning framework featuring: (1) a novel cross-modal joint learning mechanism that maximizes mutual information to bridge world knowledge and instance-level visual semantics; (2) free-form, open-vocabulary action generation; and (3) co-training of action decoding and textual reasoning. Evaluated on three benchmarks, our method achieves state-of-the-art performance, simultaneously attaining high closed-set accuracy and strong open-vocabulary robustness. It significantly improves planning accuracy, flexibility, and generalization across diverse tasks and unseen action vocabularies.
Current large language models (LLMs) rely solely on textual instructions for tool invocation, limiting their ability to accurately infer users’ underlying intentions—particularly under modality ambiguity and when multiple functionally equivalent tools are available. To address this, we propose MM-ToolLLM, the first multimodal LLM explicitly designed for tool calling, integrating vision (ViT) and audio (Whisper) encoders into a tool agent to enable cross-modal intent understanding and robust tool matching. We introduce the first multimodal tool-instruction dataset featuring multiple candidate tools per query, and adopt a joint training strategy combining instruction tuning with multimodal alignment. Experiments demonstrate significant improvements in tool recommendation accuracy, effective resolution of ambiguous queries, and support for multi-solution recommendations among functionally equivalent tools. The code and dataset are publicly released.
This study investigates whether incorporating program flowcharts into multimodal large language models can enhance code generation performance. Focusing on AtCoder programming problems, we automatically generate flowcharts at varying levels of detail and jointly input them with textual problem descriptions into GPT-4o for code synthesis. We present the first systematic evaluation of how flowchart granularity affects model performance and compare this approach against few-shot learning baselines. Experimental results demonstrate that integrating flowcharts improves code generation accuracy by up to 10%, with more detailed flowcharts yielding consistently stronger gains. Notably, one-shot learning combined with flowcharts exhibits greater stability than two-shot learning under the same conditions.
This work proposes a multimodal interaction framework that overcomes the limitations of traditional human-robot interaction, which often relies on predefined commands and lacks naturalness and expressiveness. For the first time, a large language model is leveraged to fuse semantic speech, deictic gestures, and musical beat cues through context-aware reasoning, generating coherent and expressive motion sequences for a quadruped robot. The system integrates modules for speech transcription, gesture recognition, and beat detection, employing structured prompt templates to guide the large language model and executing the resulting action queue in real time via ROS. Experimental results demonstrate that this approach significantly enhances the naturalness, flexibility, and creativity of human-robot interaction.
Existing procedural material generation methods merely replicate node graph structures without capturing the underlying design logic employed by experts, often yielding suboptimal results. This work proposes a process-driven generation paradigm that, for the first time, treats expert creation processes as first-class representations. By automatically analyzing tutorial videos, the approach extracts textualized process trajectories that encode design steps, parameter settings, and intent. Leveraging pretrained large language models, it constructs a ProcessSynthesizer and a Compiler to generate user-aligned trajectories and compile them into executable Blender material graphs. Expert evaluations demonstrate that the generated materials better reflect professional design strategies and require fewer edits, while a user study with 150 participants confirms significant improvements over existing systems in both output quality and editing efficiency.