Score
Designs and implements algorithms and systems that detect and infer when and where a human hand makes contact with an object or surface, including identifying contact frames and per-frame hand-object contact events. Produces binary or multi-state and often probabilistic contact labels and timing estimates that serve as high-level cues for sensor fusion or downstream analysis.
Accurately detecting the precise temporal moments of hand–object contact in first-person videos is challenging due to subtle motions and frequent occlusions. This work proposes the Hand-informed Context Enhanced (HiCE) module, which integrates spatiotemporal features from both hand regions and their surrounding context, augmented with a cross-attention mechanism to model latent contact patterns. To further refine temporal discrimination, the authors introduce a grasp-aware loss function and a soft-label training strategy. Evaluated on TouchMoment—a newly curated large-scale dataset comprising 8,456 annotated contact moments—the proposed method achieves a 16.91% improvement in average precision over existing event localization approaches under a two-frame tolerance metric, substantially advancing the accuracy of contact moment detection.
This work addresses the critical problem of hand-object contact state recognition in dexterous grasping. We propose a lightweight, geometry- and physics-driven contact detection method that leverages hand-object relative pose estimation and interpretable grasp quality indicators—including Grasp Quality Index (GQI) and minimum singular value of the grasp wrench matrix. Unlike existing approaches relying on dense visual annotations or image-based features, ours is the first to directly employ physically grounded grasp quality metrics for binary contact classification. By modeling geometric pose relationships and computing these metrics with adaptive thresholds, our method achieves interpretable, annotation-free contact inference without requiring RGB/RGB-D inputs or large-scale labeled data. This design significantly enhances physical plausibility and cross-scene generalizability. Evaluated on the DexYCB benchmark, the method achieves 89.7% contact detection accuracy, demonstrating both effectiveness and robustness under diverse object geometries and grasp configurations.
Existing contact identification methods in physical human–robot collaboration (pHRC) are largely limited to binary soft/hard object classification, hindering fine-grained, safety-critical interaction. Method: We propose an ontology-aware multi-class contact recognition framework distinguishing three categories—human, soft object, and hard object—using time-series force and position data from a Franka Emika Panda manipulator. We design and comparatively evaluate three end-to-end temporal architectures—LSTM, GRU, and Transformer—and rigorously assess the impact of sliding-window preprocessing. Contribution/Results: The optimized Transformer achieves 91.11% accuracy in real-time testing, substantially outperforming conventional binary classification. To our knowledge, this is the first work enabling ontology-aware, three-way contact semantic parsing in pHRC, establishing a deployable foundation for adaptive, safety-aware control in dynamic environments.
This work addresses pose-related artifacts (PRAs)—spurious signals unrelated to contact forces—that arise in flexible tactile gloves during hand posture changes, leading to false or delayed contact detection at low-force regimes and elevating the minimum detectable force (MDF). The study presents the first systematic characterization of the relationship between hand pose and these artifacts and introduces a hardware-agnostic, glove-independent pose-aware residual correction framework. By incorporating hand pose information, the method employs a dedicated residual prediction branch to explicitly compensate tactile signals. Evaluated across three distinct glove designs and fifteen users, the approach consistently reduces MDF by 10.4%, 12.2%, and 18.3%, respectively, while improving all relevant performance metrics, thereby demonstrating its generalizability and effectiveness.
This work addresses the challenge robots face in dynamically shared environments due to the absence of full-body, morphology-adaptive tactile and proximity sensing, which hinders early contact anticipation. To overcome this limitation, the authors propose GenTact-Prox—an entirely 3D-printed, modular, and programmable artificial skin capable of conforming to arbitrary robot morphologies while integrating both tactile and proximity perception. They further introduce a data-driven framework that constructs a “peripersonal space” representation for real-time pre-contact prediction. Deployed with five units on a Franka Research 3 robot, the system detects nearby objects at distances up to 18 cm, substantially enhancing safe interaction capabilities. This study presents the first demonstration of a full-body, morphology-adaptive multimodal perceptual skin, establishing a new paradigm for robotic environmental awareness.
This study addresses the challenge of accurately perceiving in-hand object contact states for robotic manipulation without dedicated tactile sensors. Inspired by human multimodal perception, the work proposes the first contact-sensing-free framework that fuses RGB-D visual inputs with proprioceptive data to generate tactile-like binary contact signals for contact state estimation. The approach employs a Transformer-based multimodal fusion architecture trained end-to-end to jointly process image and joint state information. Experimental results demonstrate that the model performs effectively in both simulated and real-world environments, generalizes to unseen objects, and successfully supports reinforcement learning tasks such as in-hand object reorientation. The method offers a low-cost, highly generalizable solution for contact-rich manipulation without requiring specialized tactile hardware.
This work addresses the challenge of capturing high-fidelity manipulation demonstrations rich in contact information while preserving human dexterity. The authors propose a novel tactile glove that, for the first time, integrates an anatomically aligned 22-degree-of-freedom joint structure, explicit contact geometry, and a high-resolution piezoresistive tactile array with 2048 sensing elements into a single wearable system. By covering the fingers, thumb, and palm with 16 rigid functional surfaces, the glove simultaneously records joint kinematics and tactile signals at 120 Hz, enabling synchronized, contact-aware capture of dexterous hand motions. This approach provides high-quality, contact-grounded demonstration data essential for advancing dexterous robotic learning.
This work addresses the challenge of accurate contact torque estimation in whole-body physical human–robot interaction, where performance is often degraded by friction disturbances and ambiguities in sensor data. The authors propose a multimodal approach that fuses pneumatic tactile skin with motor current-based proprioception to implicitly disentangle external contact forces from residual friction effects. By leveraging tactile cues and employing a temporal convolutional network (TCN) to model the hysteresis inherent in stick–slip transitions, the method achieves high-fidelity, smooth multi-axis contact force reconstruction from initial contact without requiring explicit friction identification. Experimental validation on a tactile-skin-integrated robotic arm demonstrates substantial improvements over unimodal baselines, exhibiting enhanced sensitivity and responsiveness under both static and dynamic contact conditions, while simultaneously enabling reliable force estimation and kinesthetic teaching.
This work addresses the challenge of robots failing in fine manipulation tasks due to their inability to perceive transient, weak external contacts. The authors propose TECDAR, a method that integrates a miniature 6D inertial measurement unit at the gripper tip to enable high-speed dynamic tactile sensing. By fusing this tactile data with robot pose information, TECDAR achieves real-time contact detection and localization. Leveraging an ultra-compact 6D tactile sensor operating at 7 kHz, the approach realizes sub-millisecond transient contact detection and ranging while requiring only low data bandwidth, enabling millisecond-scale trajectory correction and millimeter-level positioning accuracy. Experiments demonstrate that TECDAR achieves an average localization accuracy of 7 mm within 180 ms for both point and line contact tasks, supporting precise, purely tactile-driven manipulation and environmental perception.
This work addresses the limitations of existing RGB-based hand detection in dynamic scenarios—such as frame-rate constraints, motion blur, and poor performance under low-light conditions—and the scarcity of annotated data hindering direct application of event cameras to object detection. To bridge this gap, the authors present the first publicly available multimodal hand detection dataset from a first-person perspective. Built upon the EgoHands RGB dataset, it leverages the v2e simulator to synthesize high-temporal-resolution event streams and employs fine-tuned YOLOv8 to generate RGB bounding boxes, which are then temporally interpolated to produce aligned event-domain labels. The dataset further supports diverse lighting and scale conditions. Experiments demonstrate that multimodal detection methods trained on this dataset achieve state-of-the-art performance, validating the efficacy of synthetic event data in enhancing robustness and optimizing the bandwidth–latency trade-off.