vision-language procedural reasoning

Designs, builds, and evaluates models and pipelines that combine visual inputs and natural language (multimodal LLMs) to infer, generate, or align step-by-step procedures and action sequences from observations and instructions. This includes mapping visual observations to procedural steps, inferring high-level task contexts and progression, and producing structured procedural outputs or contextual signals for downstream modules.

vision-languageproceduralreasoning

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.42
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This work evaluates the capability of current open-source multimodal large language models (MLMs) to provide real-time assistance in technical procedural tasks, with a focus on their limitations in understanding assembly instructions and aligning textual manuals with corresponding video actions. To this end, we introduce the first fine-grained Manual-to-Action Dataset (M2AD), which precisely aligns furniture assembly videos with their respective instruction steps. Leveraging M2AD, we systematically assess MLMs on key competencies including procedural comprehension, step tracking, and cross-modal referencing between text and images. Our findings reveal that, despite some models demonstrating preliminary procedural reasoning abilities, architectural and hardware constraints hinder their effective processing of multiple image inputs and complex interleaved text–image reasoning, thereby exposing significant gaps in their applicability to real-world technical assistance scenarios.

Assembly VideosInstruction ManualsMultimodal Large Language Models

Existing multimodal large language models (MLLMs) employ only unidirectional mapping of visual information into the language space, failing to effectively leverage visual knowledge to enhance holistic reasoning. This work introduces Vision-Augmented Large Language Models (VA-LLMs), breaking from this paradigm by enabling LLMs to actively store, share, and retrieve visual knowledge. Our contributions are threefold: (1) Modular Visual Memory (MVM), a structured, long-term storage mechanism for visual knowledge; (2) Soft Multimodal Mixture of Experts (MoME), a dynamic architecture that coordinates vision and language experts during autoregressive generation; and (3) a cross-modal knowledge injection and collaborative reasoning framework. Experiments demonstrate substantial improvements in physical commonsense understanding and spatial reasoning, achieving state-of-the-art performance across multiple multimodal benchmarks, including MMMU, ScienceQA, and POPE.

Enhances LLMs with multimodal knowledge storage and sharingIntroduces Modular Visual Memory for open-world visual informationUses Mixture of Multimodal Experts to improve reasoning capabilities

Multimodal Large Language Models (MLLMs) combine the natural language understanding and generation capabilities of LLMs with perception skills in modalities such as image and audio, representing a key advancement in contemporary AI. This chapter presents the main fundamentals of MLLMs and emblematic models. Practical techniques for preprocessing, prompt engineering, and building multimodal pipelines with LangChain and LangGraph are also explored. For further practical study, supplementary material is publicly available online: https://github.com/neemiasbsilva/MLLMs-Teoria-e-Pratica. Finally, the chapter discusses the challenges and highlights promising trends.

AI integrationMultimodal Large Language Modelsmultimodal perception

PlanLLM: Video Procedure Planning with Refinable Large Language Models

Dec 26, 2024
DY
Dejie Yang
🏛️ Peking University

Existing video procedural planning methods rely on LLMs to generate fixed action sequences, suffering from poor generalization and limited adaptability to novel tasks or open-vocabulary actions. This work addresses video program planning for embodied intelligence—i.e., inferring executable action sequences from given start and goal frames. We propose an LLM-enhanced planning framework featuring: (1) a novel cross-modal joint learning mechanism that maximizes mutual information to bridge world knowledge and instance-level visual semantics; (2) free-form, open-vocabulary action generation; and (3) co-training of action decoding and textual reasoning. Evaluated on three benchmarks, our method achieves state-of-the-art performance, simultaneously attaining high closed-set accuracy and strong open-vocabulary robustness. It significantly improves planning accuracy, flexibility, and generalization across diverse tasks and unseen action vocabularies.

Adaptability to New SituationsLarge Language Models (LLMs)Video Process Planning

MLLM-Tool: A Multimodal Large Language Model for Tool Agent Learning

Jan 19, 2024
CW
Chenyu Wang
🏛️ ShanghaiTech University | Meituan | UniDT

Current large language models (LLMs) rely solely on textual instructions for tool invocation, limiting their ability to accurately infer users’ underlying intentions—particularly under modality ambiguity and when multiple functionally equivalent tools are available. To address this, we propose MM-ToolLLM, the first multimodal LLM explicitly designed for tool calling, integrating vision (ViT) and audio (Whisper) encoders into a tool agent to enable cross-modal intent understanding and robust tool matching. We introduce the first multimodal tool-instruction dataset featuring multiple candidate tools per query, and adopt a joint training strategy combining instruction tuning with multimodal alignment. Experiments demonstrate significant improvements in tool recommendation accuracy, effective resolution of ambiguous queries, and support for multi-solution recommendations among functionally equivalent tools. The code and dataset are publicly released.

Address ambiguity in user intentions with visual/auditory dataEnable multi-modal input for accurate tool selectionEnhance LLMs' tool perception beyond text queries

Latest Papers

What's happening recently
View more

This study investigates whether incorporating program flowcharts into multimodal large language models can enhance code generation performance. Focusing on AtCoder programming problems, we automatically generate flowcharts at varying levels of detail and jointly input them with textual problem descriptions into GPT-4o for code synthesis. We present the first systematic evaluation of how flowchart granularity affects model performance and compare this approach against few-shot learning baselines. Experimental results demonstrate that integrating flowcharts improves code generation accuracy by up to 10%, with more detailed flowcharts yielding consistently stronger gains. Notably, one-shot learning combined with flowcharts exhibits greater stability than two-shot learning under the same conditions.

code generationflowchartsmultimodal LLMs

This work proposes a multimodal interaction framework that overcomes the limitations of traditional human-robot interaction, which often relies on predefined commands and lacks naturalness and expressiveness. For the first time, a large language model is leveraged to fuse semantic speech, deictic gestures, and musical beat cues through context-aware reasoning, generating coherent and expressive motion sequences for a quadruped robot. The system integrates modules for speech transcription, gesture recognition, and beat detection, employing structured prompt templates to guide the large language model and executing the resulting action queue in real time via ROS. Experimental results demonstrate that this approach significantly enhances the naturalness, flexibility, and creativity of human-robot interaction.

action synthesisexpressive roboticshuman-robot interaction

Existing procedural material generation methods merely replicate node graph structures without capturing the underlying design logic employed by experts, often yielding suboptimal results. This work proposes a process-driven generation paradigm that, for the first time, treats expert creation processes as first-class representations. By automatically analyzing tutorial videos, the approach extracts textualized process trajectories that encode design steps, parameter settings, and intent. Leveraging pretrained large language models, it constructs a ProcessSynthesizer and a Compiler to generate user-aligned trajectories and compile them into executable Blender material graphs. Expert evaluations demonstrate that the generated materials better reflect professional design strategies and require fewer edits, while a user study with 150 participants confirms significant improvements over existing systems in both output quality and editing efficiency.

design intentexpert demonstrationsmaterial authoring

Hot Scholars

YX

Yijia Xiao

University of California, Los Angeles
AI for FinanceAgentsAI for ScienceMultimodal LLM
HM

Hai-Ming Xu

TikTok
Machine LearningComputer Vision
KL

Kan Li

Huazhong University of Science and Technology
3D AssemblyStretchable ElectronicsMetamaterials
FL

Feng Li

Technical University of Munich; Ph.D. Student
JL

Jianzhuang Liu

Shenzhen Institutes of Advanced Technology, University of Chinese Academy of Sciences
Computer VisionImage ProcessingAIGCMachine Learning