Score
Designs and trains control policies that are explicitly conditioned on a selected grasp or grasp parameters, converting candidate grasps into complete grasp–move–actuate sequences for object manipulation. Builds and evaluates reinforcement‑learning algorithms, reward structures, and state/observation encodings to refine policies, optimize grasp stability and motion trajectories, and support transfer from simulation to real hardware.
Dexterous manipulation by anthropomorphic robotic hands remains a core challenge in robotics, hindered by data scarcity, poor skill generalization, and the sim-to-real gap. Method: This paper proposes a unified framework integrating simulation-based environment construction, human teleoperated demonstration collection, imitation learning, and deep reinforcement learning. It innovatively unifies multi-source data acquisition paradigms with hierarchical skill learning to establish the first systematic research framework for embodied dexterous manipulation. Contribution/Results: Experimental evaluation demonstrates significant improvements in manipulation robustness and cross-task transferability. The work clarifies the technical evolution trajectory and identifies key bottlenecks, while providing a scalable methodological foundation and concrete research directions for embodied intelligence–driven dexterous manipulation.
This work addresses the challenge of effective exploration in contact-rich, sparsely rewarded environments where general-purpose reinforcement learning struggles to discover complex manipulation strategies. The authors propose a Sample-Guided RL framework that leverages a differentiable model-based solver—incorporating collision, contact, and force constraints—to construct a low-dimensional manifold of feasible states and guide policy learning through targeted sampling. By integrating black-box optimization to generate open-loop trajectories and introducing state-visit bias alongside behavioral cloning loss, the method significantly enhances both training efficiency and performance of goal-conditioned policies. Evaluated on a simplified two-sphere environment and Panda robot arm tasks, the approach substantially outperforms baseline methods, achieving high success rates in reaching statically stable states and demonstrating diverse whole-body contact-aware manipulation strategies.
To address challenges in dexterous robotic manipulation—including low operational accuracy, poor generalization, and prohibitively long training times in real-world scenarios—this paper proposes a closed-loop reinforcement learning framework integrating real-time human intervention, visual perception, and online policy optimization. The framework innovatively unifies human demonstration guidance, online error correction, and a modified Proximal Policy Optimization (PPO) algorithm, enabling both reactive and predictive dual-mode control within an end-to-end real-time architecture. Evaluated on dynamic manipulation, precision assembly, and bimanual coordination tasks, the method achieves near-100% success rates with only 1–2.5 hours of training time. Compared to baseline approaches, it improves average success rate by 2× and execution speed by 1.8×, while significantly enhancing robustness and cross-task generalization capability.
Demonstrating high-quality, physically plausible trajectories for teleoperated dexterous manipulation in contact-rich environments remains challenging due to the difficulty of acquiring consistent, diverse, and kinematically feasible human demonstrations. Method: This paper proposes a model-driven trajectory generation framework. It first identifies the high-entropy, low-consistency behavior of sampling-based planners (e.g., RRT) in contact-rich settings; then introduces a three-stage pipeline—RRT initialization, MPC-based refinement, and diffusion-model-based resampling—to jointly ensure physical feasibility, consistency, and diversity. Furthermore, it develops a goal-conditioned diffusion behavioral cloning (DBC) policy. Results: The method achieves zero-shot hardware transfer on two challenging contact-rich manipulation tasks, outperforming conventional behavioral cloning and pure planning baselines in terms of success rate, robustness, and generalization—without requiring any real-world demonstration data.
Functional grasping of complex objects (e.g., tools, household items) remains challenging when target hand poses cannot be achieved in a single step due to geometric or kinematic constraints. Method: This paper proposes a dexterous pre-grasping manipulation framework for anthropomorphic hands, based on end-to-end, demonstration-free deep reinforcement learning. It employs a unified single-policy multi-category architecture, a novel dense multi-component reward function, and a dual-path grasp representation integrating explicit pose encoding with implicit functional constraints. Training leverages the PPO algorithm and high-fidelity hand dynamics modeling, completed within three hours on a single GPU. Contribution/Results: The method autonomously performs repositioning and reorientation prior to grasping, generalizes robustly to unseen instances of trained object categories, and achieves high success rates in functional grasping—without requiring expert demonstrations or category-specific fine-tuning.
Cross-platform reuse of dexterous hand grasping policies remains challenging due to hardware-specific hand morphology and control interfaces. Method: This paper proposes a hand-agnostic, two-stage unified framework: (1) predicting displacement vectors of object surface key points—decoupled from hand anatomy—and (2) mapping these predictions to target hand joint controls via a differentiable adaptation module. We introduce finger-level geometric representations to model hand-object interactions and integrate Transformers to handle structural heterogeneity across diverse dexterous hands. A hierarchical strategy design enables decoupling of perception and control, while end-to-end differentiability supports zero-shot or few-shot transfer. Results: Extensive experiments on multiple high-DOF dexterous hands and complex objects demonstrate significant improvements over state-of-the-art baselines, validating strong generalization across hand morphologies and robust cross-hardware deployability.
This work addresses the limitations of traditional dexterous manipulation controllers, which rely on strongly assumed analytical models, and end-to-end reinforcement learning approaches, which often suffer from objective conflicts and training instability. The authors propose a skill decomposition framework that integrates physical and control-theoretic priors to decouple in-hand manipulation into analytically tractable subcomponents. By embedding theoretical constraints within each subcomponent to guide learning, this method systematically incorporates classical control knowledge into the learning pipeline for dexterous manipulation. Evaluated across diverse objects, sensor noise levels, actuation delays, and friction conditions, the approach significantly enhances policy learning stability, sample efficiency, and generalization, enabling efficient and precise in-hand repositioning and reorientation.
该研究通过三接触点接口设计,结合几何推理、模型控制和强化学习方法,解决了灵巧抓取中的定位、接近和稳定接触问题。
该研究通过结合模型预测控制(MPC)与强化学习(RL),解决了真实世界中灵巧操作任务早期探索效率低下的问题。
This work addresses the limitations of traditional non-prehensile manipulation methods, which rely on predefined target poses and struggle in real-world scenarios with diverse object configurations and unknown goals. The authors introduce the concept of a “graspability field,” reformulating manipulation as the optimization of an object-centric scalar field that quantifies graspability. This approach enables a closed-loop policy to autonomously reconfigure objects into graspable states without requiring predefined target poses or manual termination criteria. Leveraging reinforcement learning, the method jointly trains a policy network and a graspability field predictor using synthetic grasping data, achieving end-to-end integrated manipulation and grasping control. Experiments demonstrate that the proposed strategy reliably enhances object graspability in both simulation and real-robot settings, with predicted graspability distances strongly correlating with actual grasp success rates.
This work addresses the challenge of efficiently adapting general-purpose imitation policies to novel task objectives and constraints while maintaining data efficiency and deployment robustness. The authors propose an instruction-conditioned policy optimization framework that integrates imitation learning with reinforcement learning, leveraging natural language task descriptions to automatically generate reward functions. For the first time, this approach combines human feedback on intermediate trajectories with a Eureka-style reward generation mechanism to enable personalized policy refinement. Evaluated on simulated pick-and-place tasks, the method significantly outperforms feedback-free baselines, achieving enhanced robustness with reduced computational overhead and enabling efficient reuse of general policies across diverse task configurations.