🤖 AI Summary
This study addresses the limited generalization and low success rates of robotic manipulation policies under constrained demonstration budgets, where pursuing visual diversity often degrades performance. To overcome this, we propose ComManip, a paradigm that defines a "comfortable manipulation zone" and specializes the policy within this familiar visual space. During inference, the mobile base repositions the target into this zone before executing the fixed policy, thereby circumventing the conventional data diversity bottleneck. Experiments based on mainstream Vision-Language-Action (VLA) architectures, including ACT and RDT, demonstrate that ComManip improves success rates by over 20 percentage points in large-workspace, multi-task scenarios. This approach significantly enhances both data efficiency and manipulation robustness.
📝 Abstract
Training robot manipulation policies relies on costly robot demonstrations, making large-scale data collection impractical. Meanwhile, to improve policy generalization, existing approaches seek greater diversity in visual observations by varying object placements, viewpoints, and robot configurations during data collection. However, under a limited demonstration budget, this strategy forces the policy to model diverse visual observations, providing insufficient supervision to learn reliable observation-action correspondences under similar local conditions. Our study reveals that policies trained under this strategy achieve lower task success rates than those trained within a compact, visually and kinematically stable region. We refer to these stable regions as comfortable manipulation regions. To exploit this finding, we propose ComManip, a learning paradigm that specializes manipulation policies to comfortable manipulation regions. During inference, ComManip repositions the mobile base until the detected target center enters the familiar image-space range estimated from comfortable-region demonstrations. It then executes the same manipulation policy, enabling effective manipulation across diverse target locations. We conduct extensive experiments across multiple manipulation tasks, demonstration budgets, and policy families including ACT, $π_{0.5}$, RDT, OpenVLA-OFT, and SmolVLA. The results demonstrate that ComManip improves task success by roughly 20 percentage points or more across different policy architectures in large workspaces under limited demonstration budgets, suggesting that specializing manipulation policies to comfortable regions provides a more data-efficient learning paradigm for manipulation.