🤖 AI Summary
This study addresses the challenge in active visual manipulation where camera motion frequently causes the loss of critical cues and insufficient observational information. To mitigate this, we propose a unified world-action model that jointly learns observation and manipulation policies through an evidence-aware mechanism, effectively balancing viewpoint preservation with novel view acquisition. Furthermore, we introduce a pioneering training-time inversion technique to constrain frozen video priors, synergized with future video prediction for co-training. This enables the generation of bimanual and pan-tilt actions without requiring test-time inversion or candidate ranking. We also construct the RoboTwin-AV benchmark dataset. Experimental results demonstrate that our approach improves success rates on out-of-distribution tasks by 17.0% and surpasses baselines by 20.0 points on RoboTwin-AV, while achieving significantly superior performance over Fast-WAM in real-world kitchen scenarios.
📝 Abstract
Active vision manipulation requires a policy to control both its camera and its end-effectors, yet camera motion determines which evidence remains visible within finite observation windows. Acquiring a new view can displace task-critical cues, while retaining a view forgoes potentially useful observations. We formulate this as an evidence-aware retain--acquire problem and present ActiveWAM, a unified world--action model that learns observation and manipulation jointly. To this end, we propose training-time inversion which constrains a frozen video prior by task-bearing source evidence and visible temporal changes, eliminating the need for test-time inversion or candidate ranking. At deployment, the policy generates bimanual and pan/tilt actions-including stay and reacquisition behaviors-from view-aware history, and updates context from newly measured RGB observations. Future-video prediction serves as a co-training signal, while action generation requires neither future-video decoding nor optimal viewpoint annotations. We introduce RoboTwin-AV, a 50-task benchmark with executable pan/tilt control and automatically generated demonstrations. ActiveWAM improves TAVIS out-of-distribution success by up to 17.0 percentage points over the strongest baselines, achieves 20.0 additional points over Fast-WAM on RoboTwin-AV, and outperforms it by 26.7 points on real-world physical kitchen tasks.