What to Attend, What to Keep: Skill-Conditioned Visuotactile Representation with Progress-Guided Event Memory

πŸ“… 2026-09-29
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the limitations of visuotactile fusion in robotic manipulation, specifically its lack of skill adaptability and reliance on fixed temporal contexts. To overcome these challenges, we propose a skill-conditioned multimodal representation method. By introducing sparse event memory guided by skill progress alongside short-term observation tokens, and dynamically adjusting the fusion strategy via an attention mechanism, our approach achieves adaptive multimodal perception at the primitive skill level. Experimental results demonstrate that the proposed method reduces slip detection latency by 87%, decreases twist completion latency by 67.5%, and lowers progress error in blind search tasks by 92%.
πŸ“ Abstract
Robotic manipulation integrates vision, touch, and language, whose importance shifts across stages: vision guides reaching, while touch, through its evolution over time, decides grasping, alignment, and contact. Yet existing multi-modal manipulation policies typically use fixed temporal contexts and fusion strategies, despite shifts in what each modality contributes across different skills. We study how vision and touch should be combined at the level of primitive skills, asking what each skill needs from each sensor, and propose a skill-conditioned representation in which the queried skill conditions fusion over modality-specific short-term observation tokens while attending to a sparse event memory that retains terminal observations from the last $K$ executed skills. Evaluated by skill progress estimation on three contact-rich tasks, it reduces slip-detection delay by 87% against fine-tuned SOTA progress models, twist-completion delay by 67.5% against a vision-only ablation, and progress error on a blind search task by 92% through sparse event memory. Gains concentrate exactly where completion is defined by contact or task history. More broadly, our results suggest that observation formation not only policy architecture is a central challenge in multi-modal representation. Project Website: http://what-to-attend-what-to-keep.github.io/
Problem

Research questions and friction points this paper is trying to address.

robotic manipulation
multimodal representation
visuotactile fusion
skill-conditioned policy
progress estimation
Innovation

Methods, ideas, or system contributions that make the work stand out.

skill-conditioned representation
visuotactile fusion
sparse event memory
progress-guided attention
robotic manipulation
πŸ”Ž Similar Papers
No similar papers found.
πŸ’Ό Related Jobs
No related jobs found.
A
Amir-Hossein Shahidzadeh
Computer Science Department, University of Maryland, College Park, MD, USA
S
Seungjae Lee
Computer Science Department, University of Maryland, College Park, MD, USA
E
Eadom Dessalene
Computer Science Department, University of Maryland, College Park, MD, USA
S
Shanthosh Raaj Mohanram Mageswari
Computer Science Department, University of Maryland, College Park, MD, USA
S
Soroush Etemad
Computer Science Department, University of Maryland, College Park, MD, USA
Furong Huang
Furong Huang
Associate Professor of Computer Science, University of Maryland
Trustworthy AI/MLReinforcement LearningGenerative AI
C
Cornelia FermΓΌller
Computer Science Department, University of Maryland, College Park, MD, USA
Y
Yiannis Aloimonos
Computer Science Department, University of Maryland, College Park, MD, USA