PolyUMI: Accessible Visual-Tactile-Audio Data Collection for Object Inference and Manipulation

📅 2026-09-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of tactile and auditory perception in robotic manipulation, which hinders the acquisition of fine-grained contact information. To this end, we propose VisTA, a multimodal policy framework. We first construct an open-source wireless handheld platform that synchronously collects optical tactile, contact audio, and proprioceptive data, ensuring sensing geometric consistency between demonstration and execution. Subsequently, we design a token-level multimodal Transformer architecture that deeply fuses heterogeneous spatiotemporal signals for action prediction. Experimental results demonstrate that the proposed method achieves superior performance in object property inference, slip prevention control, and contact-rich manipulation tasks, significantly outperforming existing vision-dominant policies.
📝 Abstract
Humans typically rely on vision, touch, hearing, and proprioception to perceive contact and adapt their actions during manipulation. Providing robots with comparable responsiveness therefore requires hardware that can retain and use these complementary sensory signals. Most imitation-learning systems, however, observe demonstrations primarily through vision and proprioception, limiting access to contact information that is difficult to infer visually. We present PolyUMI, an open-source platform for scalable visual--tactile--audio demonstration collection and robot deployment. Its lightweight, wireless handheld gripper records synchronized wrist-camera, optical tactile, contact-audio, and proprioceptive observations without requiring a tethered workstation. The same sensing finger can be transferred to the robot end effector, preserving the sensing geometry between demonstration collection and policy execution. To effectively use these heterogeneous observations, we further introduce VisTA, a token-level multimodal policy that integrates information across sensors and time to predict contact-aware robot actions. Experiments spanning object inference, slip control, and contact-rich manipulation show that touch and audio reveal task-relevant information beyond vision and that VisTA is competitive with or outperforms existing multimodal policies. Together, PolyUMI and VisTA provide an accessible pipeline for collecting multimodal demonstrations and learning policies that perceive physical interaction beyond vision. Project Page: https://polyumi-vista.github.io
Problem

Research questions and friction points this paper is trying to address.

multimodal sensing
imitation learning
contact information
robot manipulation
visual-tactile-audio
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multimodal Imitation Learning
Visual-Tactile-Audio Sensing
Token-level Policy
Open-source Platform
Contact-rich Manipulation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
C
Conor W. Hayes
Center for Robotics and Biosystems, Northwestern Univ., Evanston, IL, USA
R
Rickmer Krohn
Interactive Robot Perception & Learning (PEARL) Lab, TU Darmstadt, Germany; Hessian.AI; Robotics Institute Germany (RIG)
A
Aravind Ramaswami
Center for Robotics and Biosystems, Northwestern Univ., Evanston, IL, USA
A
Anunth Ramaswami
Center for Robotics and Biosystems, Northwestern Univ., Evanston, IL, USA
Nils Dengler
Nils Dengler
PhD Researcher at University of Bonn
RoboticsReinforcement LearningRobot ControlMachine LearningSemantic Mapping
K
Kevin M. Lynch
Center for Robotics and Biosystems, Northwestern Univ., Evanston, IL, USA
J. Edward Colgate
J. Edward Colgate
Professor of Mechanical Engineering, Northwestern University
HapticsRoboticsTeleoperationTelemanipulation
Georgia Chalvatzaki
Georgia Chalvatzaki
Professor for Interactive Robot Perception and Learning, Technische Universität Darmstadt
RoboticsMachine LearningReinforcement LearningRobot PerceptionHRI
Matthew L. Elwin
Matthew L. Elwin
Associate Professor of Instruction, Northwestern University
robotics