hand contact estimation

Designs and implements algorithms and systems that detect and infer when and where a human hand makes contact with an object or surface, including identifying contact frames and per-frame hand-object contact events. Produces binary or multi-state and often probabilistic contact labels and timing estimates that serve as high-level cues for sensor fusion or downstream analysis.

handcontactestimation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.28
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Accurately detecting the precise temporal moments of hand–object contact in first-person videos is challenging due to subtle motions and frequent occlusions. This work proposes the Hand-informed Context Enhanced (HiCE) module, which integrates spatiotemporal features from both hand regions and their surrounding context, augmented with a cross-attention mechanism to model latent contact patterns. To further refine temporal discrimination, the authors introduce a grasp-aware loss function and a soft-label training strategy. Evaluated on TouchMoment—a newly curated large-scale dataset comprising 8,456 annotated contact moments—the proposed method achieves a 16.91% improvement in average precision over existing event localization approaches under a two-frame tolerance metric, substantially advancing the accuracy of contact moment detection.

contact momentegocentric videoframe-level detection

Hand-Object Contact Detection using Grasp Quality Metrics

Jan 13, 2025
AC
Akansel Cosgun
🏛️ Deakin University

This work addresses the critical problem of hand-object contact state recognition in dexterous grasping. We propose a lightweight, geometry- and physics-driven contact detection method that leverages hand-object relative pose estimation and interpretable grasp quality indicators—including Grasp Quality Index (GQI) and minimum singular value of the grasp wrench matrix. Unlike existing approaches relying on dense visual annotations or image-based features, ours is the first to directly employ physically grounded grasp quality metrics for binary contact classification. By modeling geometric pose relationships and computing these metrics with adaptive thresholds, our method achieves interpretable, annotation-free contact inference without requiring RGB/RGB-D inputs or large-scale labeled data. This design significantly enhances physical plausibility and cross-scene generalizability. Evaluated on the DexYCB benchmark, the method achieves 89.7% contact detection accuracy, demonstrating both effectiveness and robustness under diverse object geometries and grasp configurations.

Contact DetectionHand-Object InteractionRobot Delivery Accuracy

Multi-Class Human/Object Detection on Robot Manipulators using Proprioceptive Sensing

Aug 04, 2025
JH
Justin Hehli
🏛️ University of Zurich | MINDLab | ZHAW

Existing contact identification methods in physical human–robot collaboration (pHRC) are largely limited to binary soft/hard object classification, hindering fine-grained, safety-critical interaction. Method: We propose an ontology-aware multi-class contact recognition framework distinguishing three categories—human, soft object, and hard object—using time-series force and position data from a Franka Emika Panda manipulator. We design and comparatively evaluate three end-to-end temporal architectures—LSTM, GRU, and Transformer—and rigorously assess the impact of sliding-window preprocessing. Contribution/Results: The optimized Transformer achieves 91.11% accuracy in real-time testing, substantially outperforming conventional binary classification. To our knowledge, this is the first work enabling ontology-aware, three-way contact semantic parsing in pHRC, establishing a deployable foundation for adaptive, safety-aware control in dynamic environments.

Detect multi-class human/object contacts for robot safetyEvaluate preprocessing strategies for time-series contact analysisImprove binary classifiers with three-class detection models

This work addresses pose-related artifacts (PRAs)—spurious signals unrelated to contact forces—that arise in flexible tactile gloves during hand posture changes, leading to false or delayed contact detection at low-force regimes and elevating the minimum detectable force (MDF). The study presents the first systematic characterization of the relationship between hand pose and these artifacts and introduces a hardware-agnostic, glove-independent pose-aware residual correction framework. By incorporating hand pose information, the method employs a dedicated residual prediction branch to explicitly compensate tactile signals. Evaluated across three distinct glove designs and fifteen users, the approach consistently reduces MDF by 10.4%, 12.2%, and 18.3%, respectively, while improving all relevant performance metrics, thereby demonstrating its generalizability and effectiveness.

hand poseminimum detectable forcepose-related artifacts

This work addresses the challenge robots face in dynamically shared environments due to the absence of full-body, morphology-adaptive tactile and proximity sensing, which hinders early contact anticipation. To overcome this limitation, the authors propose GenTact-Prox—an entirely 3D-printed, modular, and programmable artificial skin capable of conforming to arbitrary robot morphologies while integrating both tactile and proximity perception. They further introduce a data-driven framework that constructs a “peripersonal space” representation for real-time pre-contact prediction. Deployed with five units on a Franka Research 3 robot, the system detects nearby objects at distances up to 18 cm, substantially enhancing safe interaction capabilities. This study presents the first demonstration of a full-body, morphology-adaptive multimodal perceptual skin, establishing a new paradigm for robotic environmental awareness.

contact anticipationperisensory spaceproximity sensing

Latest Papers

What's happening recently
View more

This study addresses the challenge of accurately perceiving in-hand object contact states for robotic manipulation without dedicated tactile sensors. Inspired by human multimodal perception, the work proposes the first contact-sensing-free framework that fuses RGB-D visual inputs with proprioceptive data to generate tactile-like binary contact signals for contact state estimation. The approach employs a Transformer-based multimodal fusion architecture trained end-to-end to jointly process image and joint state information. Experimental results demonstrate that the model performs effectively in both simulated and real-world environments, generalizes to unseen objects, and successfully supports reinforcement learning tasks such as in-hand object reorientation. The method offers a low-cost, highly generalizable solution for contact-rich manipulation without requiring specialized tactile hardware.

contact estimationdexterous manipulationproprioception

This work addresses the challenge of capturing high-fidelity manipulation demonstrations rich in contact information while preserving human dexterity. The authors propose a novel tactile glove that, for the first time, integrates an anatomically aligned 22-degree-of-freedom joint structure, explicit contact geometry, and a high-resolution piezoresistive tactile array with 2048 sensing elements into a single wearable system. By covering the fingers, thumb, and palm with 16 rigid functional surfaces, the glove simultaneously records joint kinematics and tactile signals at 120 Hz, enabling synchronized, contact-aware capture of dexterous hand motions. This approach provides high-quality, contact-grounded demonstration data essential for advancing dexterous robotic learning.

contact capturedexterous interactionhand motion

This work addresses the challenge of accurate contact torque estimation in whole-body physical human–robot interaction, where performance is often degraded by friction disturbances and ambiguities in sensor data. The authors propose a multimodal approach that fuses pneumatic tactile skin with motor current-based proprioception to implicitly disentangle external contact forces from residual friction effects. By leveraging tactile cues and employing a temporal convolutional network (TCN) to model the hysteresis inherent in stick–slip transitions, the method achieves high-fidelity, smooth multi-axis contact force reconstruction from initial contact without requiring explicit friction identification. Experimental validation on a tactile-skin-integrated robotic arm demonstrates substantial improvements over unimodal baselines, exhibiting enhanced sensitivity and responsiveness under both static and dynamic contact conditions, while simultaneously enabling reliable force estimation and kinesthetic teaching.

contact detectioncontact wrench estimationfriction hysteresis

This work addresses the challenge of robots failing in fine manipulation tasks due to their inability to perceive transient, weak external contacts. The authors propose TECDAR, a method that integrates a miniature 6D inertial measurement unit at the gripper tip to enable high-speed dynamic tactile sensing. By fusing this tactile data with robot pose information, TECDAR achieves real-time contact detection and localization. Leveraging an ultra-compact 6D tactile sensor operating at 7 kHz, the approach realizes sub-millisecond transient contact detection and ranging while requiring only low data bandwidth, enabling millisecond-scale trajectory correction and millimeter-level positioning accuracy. Experiments demonstrate that TECDAR achieves an average localization accuracy of 7 mm within 180 ms for both point and line contact tasks, supporting precise, purely tactile-driven manipulation and environmental perception.

6D dynamic tactile sensingcontact localizationdelicate manipulation

This work addresses the limitations of existing RGB-based hand detection in dynamic scenarios—such as frame-rate constraints, motion blur, and poor performance under low-light conditions—and the scarcity of annotated data hindering direct application of event cameras to object detection. To bridge this gap, the authors present the first publicly available multimodal hand detection dataset from a first-person perspective. Built upon the EgoHands RGB dataset, it leverages the v2e simulator to synthesize high-temporal-resolution event streams and employs fine-tuned YOLOv8 to generate RGB bounding boxes, which are then temporally interpolated to produce aligned event-domain labels. The dataset further supports diverse lighting and scale conditions. Experiments demonstrate that multimodal detection methods trained on this dataset achieve state-of-the-art performance, validating the efficacy of synthetic event data in enhancing robustness and optimizing the bandwidth–latency trade-off.

event-based camerasfirst-person viewhand detection

Hot Scholars

RB

Richard Bowden

Professor of Computer Vision and Machine Learning, CVSSP, University of Surrey
Computer VisionMachine learningArtificial Intelligence
RL

Rosario Leonardi

University of Catania
Computer VisionMachine LearningEgocentric Vision
KM

Kyoung Mu Lee

Professor, Department of Electrical and Computer Engineering, Seoul National University
Computer VisionMachine LearningArtificial Intelligence
DS

Daniel Sungho Jung

Seoul National University
Computer VisionVirtual HumansDigital HumansVirtual Reality