Score
Designs and implements tactile representation learning systems that map heterogeneous contact signals from different sensors and agents (human and robot) into a shared, sensor‑agnostic latent space; this includes building modality‑specific and jointly‑trained encoders (static, temporal/motion‑aware, residual/predictive) and pretraining/transfer procedures to align modalities using paired contact data. The work also defines training objectives and architectures to capture transient versus cumulative motion correlations and temporal tactile patterns so the learned embeddings preserve object‑level contact information and can be transferred or aligned with other modalities for downstream reasoning, control, or classification.
Tactile perception in embodied intelligence is hindered by spatial sparsity and the absence of global semantic context, while research on multimodal tactile fusion lacks a unified framework. This work systematically reviews relevant literature up to Q1 2026 and introduces, for the first time, a hierarchical taxonomy encompassing data modalities—such as tactile–visual and tactile–language—and three methodological pillars: perceptual recognition, cross-modal generation, and multimodal interaction. By integrating advances in deep learning and large language models, the study comprehensively surveys multimodal datasets, core algorithms, sensing hardware, and evaluation benchmarks, thereby clarifying the field’s developmental trajectory and offering a coherent theoretical foundation and systematic reference for future research.
This work addresses the limited cross-platform transferability of existing tactile perception strategies, which are heavily dependent on specific sensor modalities. To overcome this, the study proposes a unified, sensor-agnostic tactile representation by constructing a shared latent space across three heterogeneous tactile modalities—resistive, magnetic, and vision-based. This is achieved through modality-specific encoders, pairwise contact alignment signals, and joint training. The resulting representation enables zero-shot transfer of tactile policies across sensor types. Evaluated on four contact-intensive manipulation tasks, the method significantly improves average success rates from 27.5% to 45.9%, demonstrating effective disentanglement and generalization of cross-modal tactile perception and manipulation.
Existing tactile simulators struggle to accurately replicate the complex deformations and transduction mechanisms of real sensors, limiting sim-to-real transfer performance. This work proposes a multimodal representation learning framework that maps heterogeneous tactile signals—such as simulated penetration depth and real capacitive readings—into a shared latent space using modality-specific encoders. The model is trained with self-reconstruction, cross-reconstruction, and contrastive alignment losses, enabling zero-shot transfer without requiring high-fidelity simulation of raw sensory signals. By integrating multiphysics simulation to enrich embedding informativeness and leveraging a Warp-accelerated penalty-based contact model for computational efficiency, the approach achieves a 16.7% reduction in force prediction error and a 45.8% decrease in shape reconstruction error. It further demonstrates successful zero-shot cross-modal transfer across multiple downstream tasks and includes an open-sourced, efficient tactile simulation module.
To address the representation inconsistency and poor algorithmic generalizability arising from modality heterogeneity among non-optical tactile sensors, this paper proposes a cross-sensor unified tactile representation method based on an encoder-decoder architecture. The core innovation lies in the first-ever implicit feature alignment across non-visual tactile sensors: sensor-specific encoders extract modality-exclusive features and map them into a shared latent space, while a unified decoder reconstructs the original sensor signals. Joint self-supervised training is performed using force-controlled contact sequences spanning diverse object shapes and material properties. Experiments on Xela and Contactile sensors demonstrate that the proposed method significantly reduces cross-sensor reconstruction error. Moreover, the learned latent representations generalize directly to downstream tasks—such as contact geometry estimation—without fine-tuning, markedly improving model transferability and cross-sensor generalization performance.
It remains unclear at which representational level tactile supervision should be applied to enhance performance in contact-intensive manipulation within vision-language-action policies. Through linear probing analysis, this work identifies that intermediate action-expert features are best suited for predicting future tactile signals. Building on this insight, the authors propose a lightweight latent tactile prediction mechanism that leverages tactile signals as grounding supervision to align intermediate representations with anticipated contact outcomes, thereby circumventing the need to directly model noisy raw tactile data. Evaluated on real-world contact-intensive tasks, the method significantly outperforms existing non-aligned or multi-interface tactile prediction approaches and provides the first evidence of the critical role played by intermediate action representations in tactile grounding.
This work addresses the challenge of few-shot cross-task zero-shot transfer in tactile learning. We propose UniT—the first method to learn general-purpose tactile representations from a single, simple-object tactile image. UniT leverages a VQGAN-based architecture to construct a compact, disentangled latent representation space for tactile data, requiring no task-specific annotations or additional fine-tuning. The learned representations enable direct zero-shot transfer to diverse downstream tasks, including tactile perception (e.g., in-hand 3D/6D pose estimation and tactile classification) and manipulation policy learning. Compared to existing visual and tactile representation learning approaches, UniT achieves state-of-the-art performance across multiple benchmarks. It is successfully deployed on three real-world robotic manipulation tasks, demonstrating key advantages: plug-and-play usability, extreme data efficiency (single-object training), and task-agnostic generalization.
This work addresses the challenge of learning effective visuomotor policies under data scarcity while avoiding the high cost of large-scale pretraining. It proposes the τ framework, which learns action-conditioned, high-dimensional tactile spatiotemporal representations in a latent space through JEPA-inspired future visual supervision, and fuses these with pretrained vision-language features to generate actions. Tactile supervision is introduced only during training, incurring no additional overhead at deployment. Key contributions include the first use of future visual signals to model tactile dynamics, the construction of TacAura—a multimodal, temporally aligned tactile-vision-language dataset—and a novel tactile-augmented VLA architecture that requires no extra computational cost during inference. Experiments demonstrate that τ significantly outperforms existing methods across four contact-rich manipulation tasks, exhibiting strong generalization, superior task performance, and robustness.
Existing vision-language-action models struggle to effectively integrate tactile modality for modeling contact dynamics, and the limited scale and narrow scope of current tactile datasets hinder dexterous manipulation performance. To address this, this work introduces H-Tac, a large-scale tactile-action dataset comprising 160 hours of first-person human videos, and proposes TTP, a transferable tactile pretraining framework. TTP unifies human and robotic tactile-action representation spaces, explicitly models contact dynamics, and incorporates a tactile-expert-guided future prediction mechanism, enabling cross-domain knowledge transfer from large-scale human tactile data for the first time. Experiments demonstrate that TTP significantly enhances generalization and fine-grained control in dexterous manipulation tasks on both simulated and real robots.
为解决跨传感器材料识别的鲁棒性问题,提出了一种语言引导的表示学习方法,通过触觉-语言数据集训练模型以提高准确性和泛化能力。
本文提出Tactile-JEPA,一种基于触觉传感器空间布局的自监督预训练方法,以提高分布式触觉传感器表示学习的质量,从而改善机器人操作性能。
This study addresses the challenges of significant contact response discrepancies and control information loss in tactile sim-to-real transfer by proposing a decomposed tactile representation and control framework. The method introduces a novel factorized decomposition mechanism for tactile responses, decoupling contact dynamics into geometric, force distribution, and temporal variation components, which are independently evaluated through tailored encoding and randomization techniques. Additionally, a gating strategy is designed to preserve component-wise representations, enabling generalization under full-masking configurations without retraining. Experimental results demonstrate that the proposed framework achieves sub-millimeter contact localization, reduces force tracking error on unseen geometries to merely 1.69N, and improves success rates by 35% in real-world adversarial peg-in-hole tasks.