RoboMP-DINOv2: Prompts, Not Filters for Robust Robot Manipulation

📅 2026-09-21
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
为解决机器人操作中视觉变化适应性问题,提出RoboMP-DINOv2模型,通过将掩码作为空间提示而非过滤器,并结合MCR增强鲁棒性。
📝 Abstract
Robot manipulation policies must generalize across visual shifts while preserving scene context relevant to action. General-purpose vision encoders are not tailored to visuomotor control, while object-centric approaches often use segmentation masks as hard filters that discard potentially useful context. We propose RoboMP-DINOv2 (Robotics Mask-Prompted DINOv2), a full-scene vision encoder that treats masks as spatial prompts rather than visibility filters. It extracts dense DINOv2 features from the full observation, injects learned region-specific embeddings at masked locations, and jointly contextualizes prompted and unprompted tokens for action prediction. We further introduce masked-region color randomization (MCR) to improve appearance robustness, yielding RoboMP-DINOv2-MCR. Across seven simulated manipulation settings, RoboMP-DINOv2 achieves 60.7% success under spatial shifts and 59.7% under scene clutter, compared with 50.7% and 41.0% for a DINOv2-based Diffusion Policy. Under unseen object colors, RoboMP-DINOv2-MCR achieves 72.5% success versus 35.1% for the strongest color-randomized baseline. Additional experiments and representation analyses show improved robustness while preserving behaviorally relevant scene information. Code is available at https://github.com/han20192019/RoboMP_DINOv2.
Problem

Research questions and friction points this paper is trying to address.

Robot Manipulation
Visual Generalization
Scene Context
Vision Encoder
Mask Prompt
Innovation

Methods, ideas, or system contributions that make the work stand out.

RoboMP-DINOv2
spatial prompts
masked-region color randomization (MCR)
full-scene vision encoder
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.