VisForce: Visual Grounding of Current and Desired Forces for Goal-Conditioned Dexterous Manipulation

📅 2026-09-22
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
本文提出VisForce方法,通过视觉化当前与期望力的位置来解决灵巧手操作中力与视觉位置对应的问题,并在多种任务上验证了其有效性。
📝 Abstract
Vision-Language-Action (VLA) models have emerged as general-purpose robotic manipulation policies. However, in dexterous hand manipulation, contact forces are typically provided as separate states or force-specific representations, making it difficult to explicitly represent the spatial correspondence between force and their corresponding visual locations. In this work, we propose VisForce, which visually grounds the current and desired forces at their corresponding fingertip locations. VisForce renders current and desired visual force cues on the current wrist image and a task-specific goal image, and combines the two representations through goal-conditioned cross-attention to generate force-aware actions. We evaluate VisForce using a real UR10 robot equipped with an RH56F1 dexterous hand through force-conditioned grasping and three multi-stage manipulation tasks. In force-conditioned grasping experiments, VisForce exhibited a consistent grip-force response as the desired force increased, and achieved grasp-and-lift success rates of 70% and 80% for an egg and a toothpaste tube, respectively. It further achieved final success rates of 70%, 55%, and 40% on cup insertion/bottle pouring, tong-assisted bread transfer, and slip-modulated peg-in-hole, respectively. These results show that fingertip-aligned visual force representations can be effectively used for force-aware conditioning in VLA-based dexterous hand manipulation.
Problem

Research questions and friction points this paper is trying to address.

dexterous manipulation
force representation
visual grounding
spatial correspondence
vision-language-action
Innovation

Methods, ideas, or system contributions that make the work stand out.

VisForce
visual grounding of forces
goal-conditioned cross-attention
force-aware actions
🔎 Similar Papers