CompVLA: A Variable Compliance Vision-Language-Action Model for Contact-rich Manipulation

📅 2026-09-20
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决接触丰富操作任务中机器人运动与对外力响应的调控问题,提出CompVLA模型,通过RGB和语言输入联合预测运动和刚度矩阵,提高任务成功率。
📝 Abstract
Contact-rich manipulation, requiring robots to regulate not only motion but also how they yield to external forces, has emerged as the next frontier for Vision-Language-Action (VLA) models. However, existing VLAs output purely kinematic commands, degrading performance on real-world contact-rich tasks. In this paper, we introduce CompVLA, a unified VLA framework that jointly predicts motion and stiffness matrix from RGB and language inputs. Our approach augments the conventional architecture with a dedicated Compliance Expert, which outputs time-varying stiffness and virtual displacement profiles executed via geometric impedance control. We demonstrate that CompVLA achieves the highest average success rate across diverse contact-rich tasks, outperforming both vanilla and compliance-aware VLA baselines, with ablations confirming each component is essential.
Problem

Research questions and friction points this paper is trying to address.

Contact-rich manipulation
Vision-Language-Action model
Compliance
Innovation

Methods, ideas, or system contributions that make the work stand out.

Variable Compliance
Vision-Language-Action Model
Contact-rich Manipulation
Geometric Impedance Control
💼 Related Jobs
No related jobs found.
Jongmin Kim
Jongmin Kim
Seoul National University
computer architecturehomomorphic encryption
J
Junsu Ha
Seoul National University
C
Che-Sang Park
Seoul National University
M
Minchang Song
Seoul National University
H
Hyeokju Jeong
Seoul National University
H
Himchan Hwang
Seoul National University
Jianlong Fu
Jianlong Fu
Microsoft Research
Multimedia AnalysisComputer VisionRobot Learning
F
Frank C. Park
Seoul National University