TADreamer: Zero-Shot Language-Guided 3D Navigation for Terrestrial-Aerial Bimodal Robots via Video Imagination

📅 2026-09-17
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
TADreamer通过视频想象和几何校准,解决了地面-空中双模机器人在零样本情况下的语言引导3D导航问题。
📝 Abstract
Language-guided navigation for terrestrial-aerial bimodal robots requires selecting routes and locomotion modes that match scene context and task intent. Generated videos can represent such motion sequences, but recovering metrically consistent navigation references from them is challenging because of scale ambiguity and axis-dependent geometric distortions. We present TADreamer, a zero-shot framework that grounds video-imagined navigation in measured geometry without task-specific training or fine-tuning. A vision-language model translates onboard observations and instructions into navigation prompts, selects valid generated videos, and provides corrective feedback when regeneration is needed. The selected video is reconstructed into 3D waypoints annotated with terrestrial or aerial modes. A two-stage calibration procedure uses field-of-view constraints to initialize scale estimation, then refines axis-dependent scales, rotation, and translation by registering the reconstructed point cloud to measured geometry. The calibrated waypoints and mode labels guide a planner that incorporates measured geometry for robot execution. Real-world experiments demonstrate navigation across seven indoor and outdoor scenarios. With five candidates per round, usable videos are obtained within two rounds in all seven scenarios. On the calibration observations, our method reduces mean absolute depth error by 87.7% and mean absolute relative depth error by 86.3% compared with NavDreamer.
Problem

Research questions and friction points this paper is trying to address.

Language-guided navigation
bimodal robots
metrically consistent
scale ambiguity
geometric distortions
Innovation

Methods, ideas, or system contributions that make the work stand out.

Zero-Shot
Language-Guided Navigation
Video Imagination
Bimodal Robots
Calibration Procedure
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
X
Xiangyu Li
State Key Laboratory of Industrial Control Technology, Zhejiang University, Hangzhou 310027, China; Huzhou Institute of Zhejiang University, Huzhou 313000, China
T
Tiancheng Lai
State Key Laboratory of Industrial Control Technology, Zhejiang University, Hangzhou 310027, China; Huzhou Institute of Zhejiang University, Huzhou 313000, China
Xijie Huang
Xijie Huang
Hong Kong University of Science and Technology
Efficient Deep LearningModel Compression
R
Ruitian Pang
State Key Laboratory of Industrial Control Technology, Zhejiang University, Hangzhou 310027, China; Huzhou Institute of Zhejiang University, Huzhou 313000, China
S
Siqi Shen
State Key Laboratory of Industrial Control Technology, Zhejiang University, Hangzhou 310027, China; Huzhou Institute of Zhejiang University, Huzhou 313000, China
J
Juncheng Chen
State Key Laboratory of Industrial Control Technology, Zhejiang University, Hangzhou 310027, China; Huzhou Institute of Zhejiang University, Huzhou 313000, China
Z
Zaisheng Pan
Institute of Cyber-Systems and Control, Zhejiang University, Hangzhou 310027, China
C
Chao Xu
State Key Laboratory of Industrial Control Technology, Zhejiang University, Hangzhou 310027, China; Huzhou Institute of Zhejiang University, Huzhou 313000, China
F
Fei Gao
State Key Laboratory of Industrial Control Technology, Zhejiang University, Hangzhou 310027, China; Huzhou Institute of Zhejiang University, Huzhou 313000, China
Yanjun Cao
Yanjun Cao
Huzhou Institute of Zhejiang University
Multi-robot systemlocalizationUWBSLAM