🤖 AI Summary
To address the representation inconsistency and poor algorithmic generalizability arising from modality heterogeneity among non-optical tactile sensors, this paper proposes a cross-sensor unified tactile representation method based on an encoder-decoder architecture. The core innovation lies in the first-ever implicit feature alignment across non-visual tactile sensors: sensor-specific encoders extract modality-exclusive features and map them into a shared latent space, while a unified decoder reconstructs the original sensor signals. Joint self-supervised training is performed using force-controlled contact sequences spanning diverse object shapes and material properties. Experiments on Xela and Contactile sensors demonstrate that the proposed method significantly reduces cross-sensor reconstruction error. Moreover, the learned latent representations generalize directly to downstream tasks—such as contact geometry estimation—without fine-tuning, markedly improving model transferability and cross-sensor generalization performance.
📝 Abstract
Generalizable algorithms for tactile sensing remain underexplored, primarily due to the diversity of sensor modalities. Recently, many methods for cross-sensor transfer between optical (vision-based) tactile sensors have been investigated, yet little work focus on non-optical tactile sensors. To address this gap, we propose an encoder-decoder architecture to unify tactile data across non-vision-based sensors. By leveraging sensor-specific encoders, the framework creates a latent space that is sensor-agnostic, enabling cross-sensor data transfer with low errors and direct use in downstream applications. We leverage this network to unify tactile data from two commercial tactile sensors: the Xela uSkin uSPa 46 and the Contactile PapillArray. Both were mounted on a UR5e robotic arm, performing force-controlled pressing sequences against distinct object shapes (circular, square, and hexagonal prisms) and two materials (rigid PLA and flexible TPU). Another more complex unseen object was also included to investigate the model's generalization capabilities. We show that alignment in latent space can be implicitly learned from joint autoencoder training with matching contacts collected via different sensors. We further demonstrate the practical utility of our approach through contact geometry estimation, where downstream models trained on one sensor's latent representation can be directly applied to another without retraining.