VLAff: Vision-Language-Affordance Model for Unified Actionable Affordances

📅 2026-08-05
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge of transferring manipulation skills from human videos to robots, hindered by the embodiment gap between humans and machines. To bridge this gap, the authors propose VLAff, a vision–language–affordance foundation model that unifies multimodal executable affordances for the first time. By integrating first-person video, 3D scene reconstruction, and hand mesh estimation, VLAff extracts actionable affordances such as grasp poses and motion trajectories, and leverages a vision–language foundation model to enable cross-modal affordance prediction and robot action generation. The study introduces EgoAffordance, a large-scale dataset, on which the method achieves state-of-the-art performance in visual affordance prediction and successfully enables zero-shot robot manipulation and affordance-guided learning in real-world settings.
📝 Abstract
Learning manipulation skills from human videos is promising for scalable robot learning. However, the embodiment mismatch between humans and robots makes this challenging. One promising solution is to learn object-centric actionable affordances that are embodiment-agnostic. In this work, we propose a framework that leverages egocentric human videos with state-of-the-art 3D Structure-from-Motion and hand mesh reconstruction to extract actionable affordances such as visual, grasp, and trajectory affordances that explicitly encode where to interact, how to grasp, and how to move. We construct EgoAffordance, a large-scale dataset comprising 204K episodes with 5.6M visual affordances and 11.6M grasp and trajectory affordances. Building on this, we introduce VLAff, a large vision-language model-based unified foundation model that learns cross-modal correlations across all actionable affordances. Given a visual observation and instruction, VLAff generates visual affordance heatmaps, grasp poses, and trajectories, which are then converted into directly executable actions by utilizing 3D scene information. Through extensive experiments, we demonstrate that VLAff not only achieves state-of-the-art performance on visual affordance prediction, but can also be effectively applied to real robot applications such as zero-shot manipulation and affordance-guided robot learning.
Problem

Research questions and friction points this paper is trying to address.

actionable affordances
embodiment mismatch
robot manipulation
vision-language model
human video learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

vision-language model
actionable affordances
egocentric video
robot manipulation
foundation model