UniTrackPLA: Unified Panorama-Language-Action Model for Instruction-Guided Navigation and Dynamic Person Tracking

πŸ“… 2026-09-30
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the limitations of existing robotic systems that rely on forward-facing observations and decouple task policies, which lead to restricted perception and fragmented control. To overcome these challenges, this work proposes a Panoramic Language-Action Model. The model introduces a panoramic perception encoding scheme and a world-action consistency mechanism built upon a shared vision-language backbone, enabling omnidirectional environmental awareness and online replanning. By integrating instruction navigation and dynamic person tracking within a unified closed-loop control framework, it effectively bridges perception and action. Experimental results demonstrate that the proposed approach significantly improves success rates across both tasks, while real-world deployments in indoor and outdoor scenarios further validate the effectiveness of the unified control paradigm.
πŸ“ Abstract
General-purpose embodied robots should support both navigation toward language-specified destinations and dynamic person tracking under arbitrary initial target azimuths. However, existing methods typically rely on forward-facing observations and address these tasks with separate policies, limiting omnidirectional perception and unified closed-loop control. We present UniTrackPLA, a unified panorama-language-action model for instruction-guided navigation and dynamic person tracking. Its Panoramic-Aware Encoding (PAE) preserves the temporal and azimuthal structure of perspective views projected from each panorama, enabling perspective-pretrained visual encoders to process omnidirectional observations. A shared vision-language backbone grounds instructions in the panoramic context and predicts continuous robot-centric waypoint chunks for both tasks. World-Action Consistency (WAC) further predicts action-conditioned future visual states and verifies waypoint prefixes online, allowing reliable actions to be reused while triggering replanning upon inconsistency. We also introduce OmniTrackNav-Bench, comprising 5,000 simulated tracking trajectories, 10,000 simulated VLN routes, and 96 verified real-world routes, providing 919,978 waypoint-supervision instances. UniTrackPLA improves overall tracking SR from 23.50% to 35.00% and Omni-VLN SR/SPL from 13.00%/12.77% to 19.75%/19.29%. Incorporating 76 real-world routes further improves held-out EP@0.2m from 42.92% to 92.08%. Closed-loop experiments on a Go2-W robot demonstrate unified panoramic tracking and navigation across indoor and outdoor environments. The project page is at https://tw5775.github.io/UniTrackPLA.
Problem

Research questions and friction points this paper is trying to address.

embodied navigation
dynamic person tracking
panoramic perception
unified control
vision-language-action
Innovation

Methods, ideas, or system contributions that make the work stand out.

Panoramic-Aware Encoding
World-Action Consistency
Unified Panorama-Language-Action Model
Instruction-Guided Navigation
Dynamic Person Tracking
πŸ”Ž Similar Papers
No similar papers found.
πŸ’Ό Related Jobs
No related jobs found.
Pengfei Qi
Pengfei Qi
Nanyang Technological University
Optical imagingComputational imagingPolarization
H
Haoran Lin
School of Artificial Intelligence and Robotics and the National Engineering Research Center of Robot Visual Perception and Control Technology, Hunan University, Changsha, China
S
Sizhuang Chen
School of Artificial Intelligence and Robotics and the National Engineering Research Center of Robot Visual Perception and Control Technology, Hunan University, Changsha, China
K
Kai Luo
School of Artificial Intelligence and Robotics and the National Engineering Research Center of Robot Visual Perception and Control Technology, Hunan University, Changsha, China
S
Sirui Zhang
School of Artificial Intelligence and Robotics and the National Engineering Research Center of Robot Visual Perception and Control Technology, Hunan University, Changsha, China
X
Xinqi Liu
School of Artificial Intelligence and Robotics and the National Engineering Research Center of Robot Visual Perception and Control Technology, Hunan University, Changsha, China
F
Fei Cheng
School of Advanced Technology, Xi’an Jiaotong-Liverpool University, Suzhou, China
Wenrui Chen
Wenrui Chen
Hunan University
RoboticsHandsGraspingDexterous ManipulationHuman-Robot Collaboration
L
Liming Yin
Suzhou VSDeep Intelligent Technology Co., Ltd., Suzhou, China
Kailun Yang
Kailun Yang
Professor. School of Artificial Intelligence and Robotics, Hunan University (HNU); KIT; UAH; ZJU
Computer VisionComputational OpticsIntelligent VehiclesAutonomous DrivingRobotics