🤖 AI Summary
This study addresses the challenge of leveraging global information from fixed surveillance cameras for indoor robot navigation by proposing InfraVLA. The method employs a CCTV encoder to transform external static views into sequential tokens, enabling pretrained vision-language-action (VLA) models to integrate infrastructure perspectives for end-to-end decision-making. Its core contributions include the first integration of infrastructure views into a VLA framework and a two-stage fine-tuning strategy based on counterfactual data upsampling and recovery, which effectively resolves the difficulty of utilizing sparse yet critical viewpoint information. Experiments demonstrate that InfraVLA achieves success rates of 100% and over 88% on in-distribution and out-of-distribution simulation tasks, respectively. Furthermore, real-world deployment on a quadruped robot yields an 83.3% success rate, substantially outperforming the proprioception-only baseline at 29.2%.
📝 Abstract
Many indoor environments in which robots operate, such as warehouses, offices, and hospitals, already have cameras installed. They observe parts of the building that the robot cannot see from where it stands, yet navigation policies, including recent vision-language-action (VLA) models, do not use them. We propose InfraVLA, an end-to-end method that adapts a pretrained navigation VLA to such static infrastructure views: a closed-circuit television (CCTV) encoder turns each external view into tokens of the input sequence. Because the views matter only at rare decision points, fine-tuning alone did not make the policy use them in our experiments; we therefore train in two stages, on demonstrations with upsampled counterfactual data and then on recovery data. We evaluate on two simulated warehouse tasks, finding an object named in the instruction and rerouting around blocked aisles, where the deciding information is often visible only to the infrastructure cameras. Tested in distribution, InfraVLA reached a success rate of 100% on both, against 34.0% and 73.6% for a baseline without CCTV input. On out-of-distribution test sets it reached 88.2% and 88.9%. On a real quadruped fine-tuned with under 10 minutes of demonstrations, the policy reached 83.3% against 29.2% for the on-board-only baseline.