"Dear LLaVA, Please Drive": A Depth-Aware Vision-Language Agent for Closed-Loop Robotic Control

📅 2026-09-19
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出了一种基于指令学习的参数高效微调方法,使预训练的视觉-语言模型能够仅使用立体深度观测生成无碰撞路径,实现自然语言驱动的机器人控制。
📝 Abstract
Vision-language models (VLMs) provide a compelling foundation for reasoning-driven mobile navigation, offering rich contextual understanding and strong generalization from large-scale pretraining. Most existing navigation frameworks rely on imitation learning and therefore require substantial labeled trajectory data, limiting their scalability and robustness. In this work, we propose a parameter-efficient approach to fine-tune a pretrained VLM for autonomous navigation using an Imperative Learning paradigm. By optimizing against differentiable geometric cost fields rather than labeled trajectories, our model learns to generate collision-free paths exclusively from stereoscopic depth observations. We introduce a unified end-to-end navigation pipeline for natural-language-driven robotic control. This system leverages a shared VLM backbone with task-specific Low-Rank Adaptation (LoRA) modules, effectively bridging the gap from semantic target selection to low-level trajectory planning. Our approach achieves competitive Success weighted by Path Length (SPL) in unseen environments while updating less than 1% of the model's total parameters. Qualitative real-world experiments validate sim-to-real generalization and stable path planning without fine-tuning on real-world data. These results highlight a practical approach for deploying VLM-based agents on mobile robots, enabling high-level semantic navigation without the prohibitive requirement for large-scale, labeled trajectory data.
Problem

Research questions and friction points this paper is trying to address.

navigation
imitation learning
labeled trajectory data
autonomous navigation
vision-language models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Imperative Learning
Differentiable Geometric Cost Fields
Low-Rank Adaptation (LoRA)
End-to-End Navigation Pipeline
Parameter-Efficient Fine-Tuning
🔎 Similar Papers
No similar papers found.
S
Sebastian Berger
Munich University of Applied Sciences, Intelligent Vehicles Lab (IVL), 80335 Munich, Germany
K
Katharina Winter
Munich University of Applied Sciences, Intelligent Vehicles Lab (IVL), 80335 Munich, Germany
F
Fabian B. Flohr
Munich University of Applied Sciences, Intelligent Vehicles Lab (IVL), 80335 Munich, Germany