CLAP: Closed-Loop Alignment with Pressure for Precise Suction Manipulation

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of suction cup grasping mislocalization caused by visual occlusion in stacked scenarios and the inability of open-loop execution to perform real-time error correction. We propose a pressure-aware Vision-Language-Action (VLA) policy that, for the first time, incorporates measured vacuum signals as conditional inputs into the VLA model. By integrating offline joint-sequence reversal data augmentation, the pressure feedback disrupts the open-loop execution window and triggers closed-loop re-reasoning to promptly correct grasping poses. Experiments conducted on a Unitree Z1 robotic manipulator demonstrate that the proposed method achieves an average success rate of 96.67%, surpassing the baseline by over ten percentage points, with a positional error of only 15.92 mm. These results confirm the effectiveness of our approach in realizing high-precision, adaptive suction control.
📝 Abstract
Stacking and palletising demand precise placement: error left in one layer is inherited by the next, and a flat pad offers no feature to funnel a wrong pose into the right one. Top-down suction suits such dense arrangements, and suction has already been brought into vision-language-action (VLA) policies. What that work does not report, however, is a policy conditioned on a measured vacuum signal, or one that uses it to abandon an action already under way. Vision does not settle the question here, because at the moment it matters the cup and the face it holds occlude each other. We present CLAP, which makes the attachment state observable through a pressure module tapped into the vacuum line. The decoded reading replaces the suction command in the policy's proprioception, is fused with the visual features, and terminates the open-loop execution window so that the policy re-infers from a fresh observation. For data, we record goal-state disassembly on the physical robot and reverse the joint-state sequence offline, without a simulation replay. Targeted phase demonstrations, 8.3% of the training frames, cover the suction transitions and the configurations an interrupted grasp leaves behind. On a real Unitree Z1, one multi-task checkpoint reaches 96.67% average success in both colour settings, 16.67 and 10.00 points above the strongest baseline, its monochromatic four-block successes averaging 15.92 mm of error. Four ablation settings fall 5.00 to 13.33 points short. We will release code and trained weights.
Problem

Research questions and friction points this paper is trying to address.

suction manipulation
precise placement
vacuum pressure feedback
occlusion
vision-language-action policy
Innovation

Methods, ideas, or system contributions that make the work stand out.

Closed-loop alignment
Pressure feedback
Suction manipulation
Vision-language-action policy
Demonstration reversal
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Y
Yixian Zou
University of Electronic Science and Technology of China
C
Chongyang Xu
School of Aeronautics and Astronautics, Sichuan University
Y
Yuling Xin
University of Electronic Science and Technology of China
Z
Ziliang Feng
School of Aeronautics and Astronautics, Sichuan University
F
Fanman Meng
University of Electronic Science and Technology of China
Shuaicheng Liu
Shuaicheng Liu
University of Electronic Science and Technology of China
Computer VisionComputational Photography