Score
Design, build, or analyze systems that acquire and pre-process tactile sensor data — including pressure/force/torque maps and camera-like tactile images — applying filtering and denoising to produce reliable tactile signals. Fuse and integrate multiple tactile modalities and complementary sensors (e.g., vision), simulate or render tactile outputs (including vision-based and finite-element tactile simulation), and estimate contact properties such as slip, contact velocity, and object pose for downstream perception and control.
This work addresses a critical bottleneck in robotic tactile perception: the incompatibility between tactile data representations and high-level computational frameworks (e.g., computer vision). To overcome challenges arising from hardware diversity, task heterogeneity, and representational fragmentation, we propose the first systematic tactile representation framework. Specifically, we introduce six generic data structures and establish a tripartite mapping guideline linking sensor hardware, task requirements, and representation formats. Our method integrates tactile sensing modeling, structured data transformation, and cross-modal interface analysis to achieve standardized, unified representation of multi-source tactile information. We empirically characterize how representation choices affect downstream task performance, revealing underlying influence mechanisms. The framework provides reusable design principles and evidence-based selection criteria for tactile representation, thereby advancing the field toward systematic, scalable, and interoperable tactile perception systems. (149 words)
Tactile perception in embodied intelligence is hindered by spatial sparsity and the absence of global semantic context, while research on multimodal tactile fusion lacks a unified framework. This work systematically reviews relevant literature up to Q1 2026 and introduces, for the first time, a hierarchical taxonomy encompassing data modalities—such as tactile–visual and tactile–language—and three methodological pillars: perceptual recognition, cross-modal generation, and multimodal interaction. By integrating advances in deep learning and large language models, the study comprehensively surveys multimodal datasets, core algorithms, sensing hardware, and evaluation benchmarks, thereby clarifying the field’s developmental trajectory and offering a coherent theoretical foundation and systematic reference for future research.
To address the high information acquisition cost of vision-tactile sensors in contact-rich tasks and the trade-off between simulation robustness and efficiency, this paper introduces Tacchi 2.0—the first lightweight Material Point Method (MPM)-based dynamic tactile simulator integrating a pinhole camera model. It enables joint, high-fidelity generation of tactile images, marker motion images, and joint images under diverse contact modalities—including pressing, sliding, and rotating. Crucially, it embeds geometric imaging modeling directly into the physics simulation framework, reducing computational overhead by one to two orders of magnitude compared to finite element methods while enhancing cross-sensor generalizability. Experimental results demonstrate that Tacchi 2.0’s synthetic data achieve high fidelity and strong robustness across multiple vision-tactile hardware platforms, outperforming purely data-driven approaches in both accuracy and adaptability.
High acquisition cost, limited scale, and poor cross-sensor generalization of tactile data, coupled with insufficient realism and transferability of existing augmentation methods, hinder robust tactile perception. Method: We propose a low-cost, single-image-driven, two-stage controllable tactile image generation framework. For the first time, contact force and pose are explicitly incorporated as physical control signals into the generative process. A conditional diffusion model jointly optimizes a physics-constrained encoder and a tactile appearance disentanglement module, enabling physically interpretable and task-transferable data augmentation—requiring no additional hardware, only one reference image and prior force/pose information. Contribution/Results: The framework synthesizes high-fidelity, diverse tactile images. Evaluated across classification, reconstruction, and manipulation tasks, it achieves an average accuracy improvement of 12.3%, significantly enhancing model robustness and cross-device adaptability.
This work addresses critical limitations in robotic tactile perception—namely, weak sensitivity, modality scarcity, and absence of active interaction mechanisms—in human-robot cohabitation scenarios. We propose a multimodal tactile fusion architecture coupled with an active perception strategy, integrating piezoresistive, piezoelectric, capacitive, magnetic, and optical sensors. Leveraging a high-fidelity simulation platform, we generate a large-scale synthetic tactile dataset to jointly optimize sensor hardware design and perception algorithms. We systematically analyze technical bottlenecks and development pathways for tactile robots across manufacturing, healthcare, recycling, and agriculture, and introduce the first full-stack framework spanning sensing modalities, perception algorithms, and application deployment. The contributions include a scalable methodology and practical implementation guidelines for next-generation embodied intelligent tactile systems, advancing tactile robotics from passive sensing toward active physical interaction understanding.
Existing vision-based tactile sensors (e.g., DIGIT, GelSight) rely on static structured light, resulting in low imaging contrast and limited deformation sensing accuracy. To address this, we propose a novel paradigm integrating dynamic structured illumination with multi-frame image fusion: temporally programmable structured light patterns are sequentially projected, and the acquired multi-view images are registered and fused via weighted gradient-domain optimization to achieve high-fidelity surface deformation reconstruction. This approach requires no hardware modification and is the first to introduce dynamic coded illumination and image fusion into vision-based tactile sensing, opening new avenues for adaptive sensor design. Experimental results demonstrate a 42% increase in image contrast, a 35% improvement in edge sharpness, and a 51% enhancement in background discriminability—enabling robust, sub-millimeter-resolution reconstruction of surface deformations.
High-fidelity tactile modeling conflicts with real-time performance, and existing methods rely heavily on scarce real-world labeled data. Method: This paper introduces the first hydroelastic contact mechanics–based tactile sensor simulation framework, enabling continuous pressure-field modeling and sensor signal synthesis for soft–soft and soft–hard non-convex surface interactions. Implemented as an efficient, plugin-based extension in MuJoCo, it ensures physical fidelity via pressure-surface discretized integration while maintaining computational efficiency. Contribution/Results: The method achieves zero-shot sim-to-real transfer—using only synthetic data, it significantly improves real-sensor performance in object state estimation. It overcomes the long-standing accuracy–speed trade-off inherent in point-contact and finite-element approaches. The code is open-sourced and integrated into the MuJoCo ecosystem.
This work addresses the lack of compact, multimodal tactile sensors for dexterous manipulation by proposing a flexible tactile sensor that integrates sensing of slip velocity, six-axis force/torque, and pressure distribution within a single deformable structure. For the first time, these three sensing modalities are unified in one design, enabling simultaneous acquisition of slip state and rich multidimensional tactile information. Fabricated using standard PCB processes and rapid prototyping techniques, the sensor offers low cost and ease of manufacturing. Experimental results demonstrate its robust performance across diverse materials and both planar and curved contact scenarios, significantly enhancing the perceptual capability and adaptability of dexterous manipulation systems.
This work addresses the challenge of accurate contact torque estimation in whole-body physical human–robot interaction, where performance is often degraded by friction disturbances and ambiguities in sensor data. The authors propose a multimodal approach that fuses pneumatic tactile skin with motor current-based proprioception to implicitly disentangle external contact forces from residual friction effects. By leveraging tactile cues and employing a temporal convolutional network (TCN) to model the hysteresis inherent in stick–slip transitions, the method achieves high-fidelity, smooth multi-axis contact force reconstruction from initial contact without requiring explicit friction identification. Experimental validation on a tactile-skin-integrated robotic arm demonstrates substantial improvements over unimodal baselines, exhibiting enhanced sensitivity and responsiveness under both static and dynamic contact conditions, while simultaneously enabling reliable force estimation and kinesthetic teaching.
Current vision-language-action models are limited in contact-intensive manipulation tasks due to the absence of tactile perception. This work proposes TacFiLM, a method that leverages feature-wise linear modulation (FiLM) to lightweightly integrate pretrained tactile representations with intermediate visual features during fine-tuning, avoiding complex token concatenation or extensive retraining. TacFiLM significantly improves success rates, direct insertion performance, task completion efficiency, and force stability in insertion tasks. Moreover, it demonstrates strong robustness across both in-distribution and out-of-distribution scenarios, achieving efficient and generalizable multimodal tactile enhancement.
Existing optical tactile methods struggle to accurately discern subtle contact states due to their reliance on raw images or accumulated motion fields, often resulting in perceptual ambiguity. To address this limitation, this work proposes a dynamic tactile representation that integrates both instantaneous and cumulative motion correlations and, for the first time, explicitly incorporates dynamic priors of tactile motion to distinguish fine-grained contact differences. Building upon this representation, the authors design a unified multimodal fusion architecture based on a Mixture-of-Transformers, which effectively preserves the distinct characteristics of visual and tactile modalities while enabling rich cross-modal interactions. Evaluated on contact-intensive manipulation tasks, the proposed approach significantly outperforms existing tactile representations and fusion strategies, demonstrating enhanced sensitivity to minute contact variations and improved manipulation performance.