DPed-VLN: A Benchmark for Socially Compliant Vision-and-Language Navigation in Dynamic Pedestrian Environments

📅 2026-09-18
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
本文提出DPed-VLN基准,用于评估动态行人环境下的视觉-语言导航,并通过强化学习和模仿学习训练的DPet模型解决社交安全约束下的导航问题。
📝 Abstract
Vision-and-language navigation (VLN) has advanced rapidly in static indoor environments, but robots operating in human-populated spaces must ground language while responding to moving pedestrians and social-safety constraints. We present DPed-VLN, a Habitat 3.0 benchmark for dynamic-pedestrian VLN that couples 33,093 navigation episodes with paired global and prior-augmented instructions, ORCA-controlled humanoid pedestrians, socially constrained expert paths, and metrics that jointly assess navigation efficiency and social safety. DPed-VLN separates ordinary goal-oriented route guidance from prior-augmented instructions that expose dynamic-pedestrian cues for controlled analysis. To instantiate the benchmark, we introduce DPet (Dynamic Pedestrian-aware Network), a pedestrian-aware policy network trained with reinforcement learning and imitation learning. We further adapt representative state-of-the-art VLM-based navigation models, including NaVILA and StreamVLN, to DPed-VLN through LoRA fine-tuning. Experiments show that LoRA adaptation improves zero-shot VLM baselines in several success and safety metrics, especially reducing StreamVLN's collision rate. Among the evaluated methods, DPet-RL achieves the highest SR, SPL, and STL.
Problem

Research questions and friction points this paper is trying to address.

Vision-and-Language Navigation
Dynamic Pedestrian Environments
Social Safety
Navigation Efficiency
Innovation

Methods, ideas, or system contributions that make the work stand out.

DPed-VLN
socially compliant navigation
dynamic pedestrian environment
DPet
LoRA fine-tuning
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
H
Haojie Dai
Department of Control Science and Engineering, Tongji University, Shanghai, China
X
Xiangyi Wang
Department of Control Science and Engineering, Tongji University, Shanghai, China
Liuyi Wang
Liuyi Wang
Tongji University
computer visionnatural language processingartificial intelligence
K
Kai Sheng
Department of Control Science and Engineering, Tongji University, Shanghai, China
Z
Zongtao He
KEENON Robotics Co., Ltd., Shanghai, China
C
Chengju Liu
Department of Control Science and Engineering, Tongji University, Shanghai, China
Wei Ye
Wei Ye
Tongji University
data miningmachine learningrepresentation learningdeep learning
Q
Qijun Chen
Department of Control Science and Engineering, Tongji University, Shanghai, China