grasp-conditioned reinforcement learning

Designs and trains control policies that are explicitly conditioned on a selected grasp or grasp parameters, converting candidate grasps into complete grasp–move–actuate sequences for object manipulation. Builds and evaluates reinforcement‑learning algorithms, reward structures, and state/observation encodings to refine policies, optimize grasp stability and motion trajectories, and support transfer from simulation to real hardware.

grasp-conditionedreinforcementlearning

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.14
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the challenge of effective exploration in contact-rich, sparsely rewarded environments where general-purpose reinforcement learning struggles to discover complex manipulation strategies. The authors propose a Sample-Guided RL framework that leverages a differentiable model-based solver—incorporating collision, contact, and force constraints—to construct a low-dimensional manifold of feasible states and guide policy learning through targeted sampling. By integrating black-box optimization to generate open-loop trajectories and introducing state-visit bias alongside behavioral cloning loss, the method significantly enhances both training efficiency and performance of goal-conditioned policies. Evaluated on a simplified two-sphere environment and Panda robot arm tasks, the approach substantially outperforms baseline methods, achieving high success rates in reaching statically stable states and demonstrating diverse whole-body contact-aware manipulation strategies.

constrained samplingcontact-rich manipulationreinforcement learning

To address challenges in dexterous robotic manipulation—including low operational accuracy, poor generalization, and prohibitively long training times in real-world scenarios—this paper proposes a closed-loop reinforcement learning framework integrating real-time human intervention, visual perception, and online policy optimization. The framework innovatively unifies human demonstration guidance, online error correction, and a modified Proximal Policy Optimization (PPO) algorithm, enabling both reactive and predictive dual-mode control within an end-to-end real-time architecture. Evaluated on dynamic manipulation, precision assembly, and bimanual coordination tasks, the method achieves near-100% success rates with only 1–2.5 hours of training time. Compared to baseline approaches, it improves average success rate by 2× and execution speed by 1.8×, while significantly enhancing robustness and cross-task generalization capability.

Enabling autonomous robotic manipulation skills in real-world settings.Improving success rates and execution speed in dexterous manipulation tasks.Learning robust, adaptive policies for reactive and predictive control strategies.

Should We Learn Contact-Rich Manipulation Policies from Sampling-Based Planners?

Dec 12, 2024
HZ
Huaijiang Zhu
🏛️ New York University | Boston Dynamics AI Institute | Georgia Tech | Cornell University | Artificial and Natural Intelligence Toulouse Institute

Demonstrating high-quality, physically plausible trajectories for teleoperated dexterous manipulation in contact-rich environments remains challenging due to the difficulty of acquiring consistent, diverse, and kinematically feasible human demonstrations. Method: This paper proposes a model-driven trajectory generation framework. It first identifies the high-entropy, low-consistency behavior of sampling-based planners (e.g., RRT) in contact-rich settings; then introduces a three-stage pipeline—RRT initialization, MPC-based refinement, and diffusion-model-based resampling—to jointly ensure physical feasibility, consistency, and diversity. Furthermore, it develops a goal-conditioned diffusion behavioral cloning (DBC) policy. Results: The method achieves zero-shot hardware transfer on two challenging contact-rich manipulation tasks, outperforming conventional behavioral cloning and pure planning baselines in terms of success rate, robustness, and generalization—without requiring any real-world demonstration data.

Address high entropy in sampling-based planner demonstrationsEnable zero-shot transfer to hardware for complex tasksGenerate training data for contact-rich dexterous manipulation tasks

Functional grasping of complex objects (e.g., tools, household items) remains challenging when target hand poses cannot be achieved in a single step due to geometric or kinematic constraints. Method: This paper proposes a dexterous pre-grasping manipulation framework for anthropomorphic hands, based on end-to-end, demonstration-free deep reinforcement learning. It employs a unified single-policy multi-category architecture, a novel dense multi-component reward function, and a dual-path grasp representation integrating explicit pose encoding with implicit functional constraints. Training leverages the PPO algorithm and high-fidelity hand dynamics modeling, completed within three hours on a single GPU. Contribution/Results: The method autonomously performs repositioning and reorientation prior to grasping, generalizes robustly to unseen instances of trained object categories, and achieves high success rates in functional grasping—without requiring expert demonstrations or category-specific fine-tuning.

Achieving human-like functional graspsLearning dexterous pre-grasp manipulationUsing deep reinforcement learning

Learning Cross-hand Policies for High-DOF Reaching and Grasping

Apr 14, 2024
QS
Qijin She
🏛️ National University of Defense Technology | Hunan University | Academy of Military Science | Shenzhen University

Cross-platform reuse of dexterous hand grasping policies remains challenging due to hardware-specific hand morphology and control interfaces. Method: This paper proposes a hand-agnostic, two-stage unified framework: (1) predicting displacement vectors of object surface key points—decoupled from hand anatomy—and (2) mapping these predictions to target hand joint controls via a differentiable adaptation module. We introduce finger-level geometric representations to model hand-object interactions and integrate Transformers to handle structural heterogeneity across diverse dexterous hands. A hierarchical strategy design enables decoupling of perception and control, while end-to-end differentiability supports zero-shot or few-shot transfer. Results: Extensive experiments on multiple high-DOF dexterous hands and complex objects demonstrate significant improvements over state-of-the-art baselines, validating strong generalization across hand morphologies and robust cross-hardware deployability.

Adaptability to Different Hand TypesRobot LearningUniversal Gripping Skills

Latest Papers

What's happening recently
View more

This work addresses the limitations of traditional dexterous manipulation controllers, which rely on strongly assumed analytical models, and end-to-end reinforcement learning approaches, which often suffer from objective conflicts and training instability. The authors propose a skill decomposition framework that integrates physical and control-theoretic priors to decouple in-hand manipulation into analytically tractable subcomponents. By embedding theoretical constraints within each subcomponent to guide learning, this method systematically incorporates classical control knowledge into the learning pipeline for dexterous manipulation. Evaluated across diverse objects, sensor noise levels, actuation delays, and friction conditions, the approach significantly enhances policy learning stability, sample efficiency, and generalization, enabling efficient and precise in-hand repositioning and reorientation.

dexterous manipulationin-grasp repositioningreinforcement learning

This work addresses the limitations of traditional non-prehensile manipulation methods, which rely on predefined target poses and struggle in real-world scenarios with diverse object configurations and unknown goals. The authors introduce the concept of a “graspability field,” reformulating manipulation as the optimization of an object-centric scalar field that quantifies graspability. This approach enables a closed-loop policy to autonomously reconfigure objects into graspable states without requiring predefined target poses or manual termination criteria. Leveraging reinforcement learning, the method jointly trains a policy network and a graspability field predictor using synthetic grasping data, achieving end-to-end integrated manipulation and grasping control. Experiments demonstrate that the proposed strategy reliably enhances object graspability in both simulation and real-robot settings, with predicted graspability distances strongly correlating with actual grasp success rates.

grasp planninggraspabilitynon-prehensile manipulation

This work addresses the challenge of efficiently adapting general-purpose imitation policies to novel task objectives and constraints while maintaining data efficiency and deployment robustness. The authors propose an instruction-conditioned policy optimization framework that integrates imitation learning with reinforcement learning, leveraging natural language task descriptions to automatically generate reward functions. For the first time, this approach combines human feedback on intermediate trajectories with a Eureka-style reward generation mechanism to enable personalized policy refinement. Evaluated on simulated pick-and-place tasks, the method significantly outperforms feedback-free baselines, achieving enhanced robustness with reduced computational overhead and enabling efficient reuse of general policies across diverse task configurations.

Human InstructionsImitation LearningPolicy Refinement

Hot Scholars

PH

Pu Hua

IIIS, Tsinghua University
Robot Learning
YJ

Yeonwoo Jeong

Seoul National University
Machine learningCombinatorial optimization
JF

Jing Fang

Northwestern Polytechnical University
Image ProcessingDeep Learning
SR

Spandan Roy

Assistant Professor, Robotics Research Center, IIIT Hyderabad
Adaptive-robust controlSwitched systemsArtificial delay controlRobotics
HX

Huazhe Xu

Tsinghua University
Embodied AIReinforcement LearningComputer VisionDeep Learning