🤖 AI Summary
This study addresses the lack of evaluation for active visual recovery under occlusion in robotic manipulation benchmarks by constructing BAVO-Bench, a bipedal active vision benchmark, and proposing the A-FAR strategy. This strategy pioneers an occlusion recovery benchmark that systematically controls external visibility. By unifying a robot-centric 3D representation framework with knowledge distillation from pretrained 4D models, it introduces a geometry-guided mechanism that operates without future observations, enabling joint optimization of viewpoint selection and manipulation control. Experimental results demonstrate that the proposed method significantly enhances robotic robustness under both structured and random temporal occlusions while maintaining high performance in unoccluded scenarios.
📝 Abstract
Physical active vision allows robots to change their viewpoint when task-relevant observations become unreliable, yet existing manipulation benchmarks provide limited support for studying how policies recover from occlusion during execution. We introduce BAVO-Bench (Bimanual Active Vision under Occlusion), a bimanual active-vision benchmark that systematically controls external visibility through Clean, Stage Occlusion, and Random-time Occlusion conditions, enabling evaluation of both manipulation performance and active visual recovery. Building on this setting, we present A-FAR (Active Future-Aware Recovery), an active-vision policy for joint viewpoint and manipulation control. A-FAR represents moving-camera observations in a unified robot-centric 3D frame and distills relational structure together with its future evolution from a pretrained 4D model, providing the policy with future-aware geometric guidance without requiring future observations at deployment. Experiments across multiple manipulation tasks show that A-FAR improves robustness to both structured and temporally shifted occlusions while maintaining strong performance under clean observations.