TouchScale: 500 Hours of Human Vision and Touch for Visual-Tactile Learning

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses critical bottlenecks in visuo-tactile learning, including data scarcity, sensor heterogeneity, and the difficulty of decoupling scaling effects. We introduce the first 500-hour unified wearable multimodal dataset that synchronously captures RGB-D, wrist-mounted video, and dense bimanual tactile signals. A standardized protocol effectively isolates and validates the crucial performance gains attributable to data scale, while pretraining combined with policy fine-tuning achieves cross-modal temporal alignment. Experimental results demonstrate substantial improvements: zero-shot contact IoU increases from 0.134 to 0.383, real-world manipulation success rates rise from 22.5% to 57.5%, and action recognition accuracy reaches state-of-the-art levels. These findings provide foundational support for learning physical interactions in embodied intelligence.
📝 Abstract
Large-scale egocentric human interaction data is becoming an important source of physical supervision for embodied learning, yet video alone leaves the contact and pressure that characterize physical interaction unrecorded. Recent visual-tactile datasets provide this missing supervision, but their synchronized tactile data remain far smaller in volume than human video. Moreover, the largest resources often merge recordings from different sensors or annotation procedures, which makes the effect of data scale difficult to isolate. We therefore introduce TouchScale, a 500-hour dataset of contact-rich human interaction recorded with a single unified wearable setup. Its approximately 2K predefined task descriptions span everyday activities and structured manipulation, and each recording temporally aligns egocentric RGB-D video with wrist RGB video and dense full-hand bimanual tactile measurements. Compared with prior tactile data, training on the full TouchScale raises zero-shot contact IoU on data from an unseen tactile sensor from 0.134 to 0.383. Pretraining a visual encoder on TouchScale also yields the highest action recognition accuracy on three benchmarks among the compared visual-tactile datasets. Used for visual-tactile mid-training of a robot policy, TouchScale improves the average real-world success rate across four contact-rich manipulation tasks from 22.5% to 57.5%. With the sensor and collection protocol held fixed, both zero-shot tactile prediction and robot success show an overall upward trend as more TouchScale data is used. These results suggest that human visual-tactile data collected at scale with consistent sensing benefits both perception and robot manipulation. We will publicly release TouchScale, including all synchronized visual-tactile recordings and reconstructed object models, to support future research on scalable visual-tactile learning.
Problem

Research questions and friction points this paper is trying to address.

visual-tactile learning
embodied learning
tactile dataset
data scale
human interaction
Innovation

Methods, ideas, or system contributions that make the work stand out.

Visual-Tactile Learning
Egocentric Dataset
Tactile Sensing
Robot Manipulation
Embodied AI
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
D
Dayou Li
Texas A&M University
Hao Wang
Hao Wang
Google
Qianqian Yang
Qianqian Yang
Zhejiang University
Information TheoryWireless AISemantic CommunicationMachine Learning
Zihao Zhu
Zihao Zhu
The Chinese University of Hong Kong, Shenzhen
AI securityLarge language modelsAgentEmbodied AI
Haoquan Fang
Haoquan Fang
University of Washington, Allen Institute for AI
Computer VisionMachine LearningEmbodied AIRobotics
Ziyao Zeng
Ziyao Zeng
Yale University
Computer VisionMachine LearningRoboticsMultimodal Learning
Yan Han
Yan Han
PhD student, Australian National University
machine learningcomputer vision
Z
Zihan Wang
Overfit Lab
Yan Wang
Yan Wang
Senior Research Scientist, NVIDIA Research
Computer VisionMachine LearningAutonomous Vehicle
Baoru Huang
Baoru Huang
University of Liverpool; Imperial College London
RoboticsComputer visionSurgical visionImage-Guided Intervention
Dilin Wang
Dilin Wang
Facebook
Machine Learning
Kenji Shimada
Kenji Shimada
Carnegie Mellon University
RoboticsCAD/CAECV/CGAIMLBME
Yiyue Luo
Yiyue Luo
Assistant Professor, University of Washington
Intelligent TextilesDigital FabricationHCIApplied Machine Learning
Manling Li
Manling Li
Assistant Professor at Northwestern University
Natural Language ProcessingVision-LanguageEmbodied Agents
T
Teresa Lv
Sony
Mustafa Mukadam
Mustafa Mukadam
Amazon Robotics
RoboticsArtificial Intelligence
Rakesh Ranjan
Rakesh Ranjan
IIT (ISM), Dhanbad
Wireless CommunicationSignal Processing
Ruohan Zhang
Ruohan Zhang
Stanford University
RoboticsCognitive ScienceBrain-Machine InterfaceMachine LearningArt
Qi He
Qi He
Microsoft
Artificial IntelligenceData MiningInformation RetrievalWeb SearchKnowledge Graph
Changliu Liu
Changliu Liu
Associate Professor, Carnegie Mellon University
Roboticshuman-robot interactionsmotion planningoptimizationmulti-agent system
Xu Chen
Xu Chen
University of Washington
Dynamic systems and controlsadditive and advanced manufacturingagile roboticsinformation fusion
Marco Pavone
Marco Pavone
Stanford University and NVIDIA
RoboticsControl TheoryDistributed ControlIntelligent Transportation systems
B
Bangya Liu
Overfit Lab
J
Jiachen Li
Georgia Tech
Masayoshi Tomizuka
Masayoshi Tomizuka
Mechaniccal Engineering, University of California
mechanical engineeringdynamic systemscontrolmechatronics