Score
Design, build, and analyze reinforcement learning systems and training pipelines that explicitly sense, model, or incorporate contact events and contact-related signals (e.g., impact events, contact sequences, pseudo‑tactile observations, and ground compliance) into state representations, policies, reward functions, or supervision. This includes constructing contact-aware simulators and dynamics models, contact‑based reward shaping or event‑sequence supervision, and algorithms that adapt policy learning to different contact modalities (rigid vs. soft) and terrain interactions.
This study addresses a critical gap in the literature: the absence of a unified survey on robotic learning methods that integrate force and tactile perception, particularly regarding the synthesis of multimodal sensing and multi-stage system design. To bridge this gap, the paper introduces the TF-ART categorization framework—a novel, comprehensive architecture that systematically encompasses multimodal perceptual inputs, hierarchical action generation, and reactive low-level control. By integrating heterogeneous sensor encoding, multimodal perception fusion, and action refinement mechanisms, the framework elucidates the intrinsic relationships among existing approaches and clearly maps their design logic across the perception–decision–execution pipeline. This contribution provides a holistic perspective, theoretical foundation, and practical guidance for developing intelligent systems capable of rich physical interaction.
Dexterous manipulation tasks with high contact complexity—such as bottle cap opening—pose significant challenges for robotic learning due to sparse rewards, high-dimensional tactile feedback, and stringent requirements on grasp stability and multi-finger coordination. Method: We propose a tactile-aware reinforcement learning framework that integrates haptic sensing with policy optimization. Key components include a tactile-guided reward shaping mechanism explicitly modeling stable grasping and coordinated multi-finger motion, and an embedded multimodal observation space jointly encoding tactile signals as both reward components and critical state inputs. Contribution/Results: The method achieves superior data efficiency and policy robustness, converging rapidly in simulation and transferring zero-shot to a Shadow Robot dexterous hand for successful real-world cap opening. Experiments demonstrate strong generalization across diverse object geometries and dynamic disturbances. This work delivers a deployable, end-to-end solution for tactile-driven dexterous manipulation.
This work addresses the challenge of contact-intensive manipulation tasks failing during sim-to-real transfer due to discrepancies in contact dynamics. To bridge this gap, the authors propose a contact-centric real-to-sim-to-real reinforcement learning framework that jointly models task-relevant contact event sequences and local contact dynamics from real-world demonstrations. By optimizing the contact geometry of objects in simulation, the method aligns simulated state transitions with their real-world counterparts and automatically generates structured reward signals without manual task-specific reward engineering. Experimental results demonstrate that the proposed approach significantly outperforms unconstrained reinforcement learning baselines across multiple contact-intensive tasks, achieving more stable and robust high-fidelity policy transfer.
This work addresses the challenge of designing separate robotic policies for each task in multi-task settings. We propose a unified locomotion and manipulation learning framework grounded in explicit contact goal sequences—specifying contact locations, temporal ordering, and end-effectors. Methodologically, we introduce contact behavior as the central task definition, establishing a goal-conditioned reinforcement learning paradigm where contact plans serve as policy inputs, enabling end-to-end training across morphologically diverse platforms (quadrupedal, bipedal, and dual-arm robots). Our key contributions are: (1) a contact-driven, generalizable task representation that captures structural commonalities across diverse tasks; (2) a single policy capable of generalizing to heterogeneous locomotion and manipulation tasks; and (3) strong robustness and cross-scenario, cross-morphology transferability, even under zero-shot conditions. Extensive experiments demonstrate effectiveness on complex physical interaction tasks.
This work addresses the limited online adaptation capability of existing vision-language-action (VLA) models in contact-intensive manipulation tasks, which often leads to improper contact force control and inefficient retries. To overcome this, the authors propose a tactile-guided online reinforcement learning framework that integrates tactile-perception-informed reference action prediction with lightweight policy optimization. The approach further introduces an intervention-masking critic mechanism to effectively fuse human intervention signals with exploration data, thereby enhancing policy robustness. Experimental results demonstrate that the proposed method significantly improves both subtask and full-task success rates, as well as time-constrained execution efficiency, across challenging tasks such as latch manipulation, coffee cup placement, and egg grasping.
This work addresses the challenges of insufficient tactile perception and difficult sim-to-real transfer in in-hand object translation with dexterous hands. Methodologically, we propose a three-axis tactile-driven control framework enabling zero-shot sim-to-real transfer: (i) we develop the first physics-consistent tactile skin model capable of simulating 3D shear and normal forces; (ii) we design a deep reinforcement learning policy—based on Proximal Policy Optimization (PPO)—that fuses multi-dimensional tactile and proprioceptive sensing, augmented with sliding-contact modeling to enhance robustness during dynamic interactions. Contributions include: (i) achieving zero-shot sim-to-real transfer without real-world fine-tuning; (ii) experimentally validating stable in-hand translation on a physical dexterous hand across unseen objects and multiple object orientations; and (iii) demonstrating that the full three-axis tactile policy significantly outperforms unimodal baselines (shear-only, normal-only, or proprioception-only), establishing a generalizable and deployable paradigm for tactile dexterous manipulation.
This work addresses the performance limitations of vision-language-action (VLA) models in contact-intensive tasks due to their lack of fine-grained tactile perception. To overcome this, the authors propose a scalable framework that integrates tactile feedback through a realistically aligned closed-loop simulator, eliminating the need for large-scale tactile pretraining or extensive real-world exploration. The approach combines hybrid sim-to-real trajectory warmstarting, a tactile modulation mechanism, and reinforcement learning guided by validation-reward signals to enhance policy robustness and distributional consistency. Notably, the method enables zero-shot transfer to real-world settings without online fine-tuning. Evaluated on four dual-arm contact-intensive tasks, it achieves an average success rate of 72.5%, substantially outperforming the 50.0% baseline.
This work addresses the loss of rich contact details in sim-to-real transfer due to oversimplified tactile representations by proposing a physics-informed center-of-pressure (CoP) tactile representation. This approach preserves dense contact information while enhancing transfer robustness and enables sensor orientation calibration without requiring ground-truth force measurements through differentiable dynamics. Integrated with reinforcement learning and multi-fingered dexterous hand control, the system achieves zero-shot sim-to-real transfer on vision-deprived tasks such as plug insertion and ball balancing. It significantly outperforms binary contact and raw tactile baselines, with the learned policy implicitly encoding physical properties of objects, including mass.
This work addresses the significant performance degradation of existing physics-informed goal-conditioned reinforcement learning methods in contact-rich manipulation tasks, where hybrid contact dynamics induce non-smooth value landscapes. To overcome this challenge, the authors propose a contact-aware hierarchical physics-informed reinforcement learning framework that reliably extends Pi-GCRL to such settings for the first time. The approach integrates optimal control–inspired inductive biases, hybrid system modeling, and a hierarchical policy architecture to selectively impose physical priors during goal-conditioned value learning. This design effectively mitigates the adverse effects of discontinuous dynamics while preserving sample efficiency. Experimental results demonstrate that the proposed method substantially improves both generalization to arbitrary goals and task success rates in environments characterized by frequent and complex contacts.
End-to-end learning approaches often suffer from poor generalization and insufficient robustness in high-precision, contact-rich manipulation tasks. To address this limitation, this work proposes a hybrid control framework that leverages offline deep reinforcement learning only during task-critical phases, where data are autonomously and densely collected via an automated mechanism, while relying on conventional motion planning for all other stages. Notably, the method requires neither teleoperation nor online policy updates, substantially improving both learning efficiency and robustness. Evaluated on four real-world tasks, the system achieves an average success rate of 96% using only 2–2.5 hours of autonomously collected data—significantly outperforming the strongest baseline at 55%—and maintains strong performance even in out-of-distribution scenarios.