VLMs Can Describe, But Not Measure: Object-Centric Scene Understanding for Robotic Manipulation

📅 2026-09-23
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
本文针对机器人在未知环境中的操作问题,提出了一种基于视觉-语言模型的模块化感知框架,通过结合RGB-D数据和深度信息来提高场景理解和定位准确性。
📝 Abstract
Robotic operation in previously unseen environments requires both semantic understanding and reliable metric information. While vision--language models (VLMs) provide strong semantic capabilities, their geometric estimates remain less reliable. In this paper, we propose a VLM-driven, modular perception framework for scene understanding using off-the-shelf approaches. Starting from a single RGB-D observation, the scene is segmented into object-level regions, annotated by a VLM, and grounded with depth information to construct a task-independent object-centric representation. Experiments on 151 tabletop scenes show that the proposed decomposition preserves strong semantic performance while substantially improving localization and depth estimation over direct VLM inference. The resulting representation is also integrated with a task-planning framework for robotic execution.
Problem

Research questions and friction points this paper is trying to address.

Robotic Operation
Semantic Understanding
Metric Information
Vision-Language Models
Geometric Estimates
Innovation

Methods, ideas, or system contributions that make the work stand out.

VLM-driven
modular perception framework
object-centric representation
semantic understanding
depth estimation
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
E
Enrico Saccon
Department of Information Engineering and Computer Science, University of Trento, Trento, Italy
T
Tommaso Faraci
Department of Information Engineering and Computer Science, University of Trento, Trento, Italy
I
Iñigo De La Ossa Zarzuelo
Polytechnic School, Mondragón University, Mondragón, Spain
L
Luigi Palopoli
Department of Information Engineering and Computer Science, University of Trento, Trento, Italy
Marco Roveri
Marco Roveri
University of Trento - Department of Information Engineering and Computer Science
Formal MethodsArtificial IntelligenceComputer Science
Matteo Saveriano
Matteo Saveriano
Associate Professor, University of Trento
RoboticsMachine LearningAI