Score
Designs and implements algorithms and data structures that detect or segment objects, sample and represent sets of feasible kinematic grasp poses, and filter or rank those grasp candidates by feasibility or quality. Builds interfaces and query operators to export and serve candidate grasps to downstream planners and controllers.
To address the lack of visual debugging support in the GRASP programming system, this paper designs and implements a modular Visual Stepper for Scheme, built upon GRASP’s extensible architecture. The stepper employs visual evaluation techniques to enable interactive program execution tracing and real-time state visualization. Its key contribution is a lightweight, reusable debugging component integration mechanism—demonstrating both the flexibility and practicality of GRASP’s extension framework and establishing a generalizable design paradigm for debugging tools in pedagogical programming environments. The prototype has been successfully integrated into GRASP and deployed in undergraduate programming instruction. Empirical classroom use shows significant improvements in novice learners’ comprehension of program evaluation semantics and debugging efficiency. This work lays a foundational technical basis for developing integrated educational toolchains supporting computational thinking and program reasoning.
Complex dexterous manipulation tasks—such as axial rotation—involve time-varying dynamics, necessitating multi-stage grasping and repeated regrasping, which challenge conventional grasp planning frameworks. Method: This paper proposes a point-cloud-driven closed-loop grasp planning method. It formalizes complex manipulation as a sequence of constant screw motions and automatically determines regrasp timing and count via intersection analysis of graspable regions. The method integrates point-cloud-based geometric modeling, screw-motion-constrained trajectory optimization, and grasp quality evaluation to synthesize continuous grasp–regrasp sequences under path constraints. Results: Evaluated on real-world RGB-D point cloud data, the approach robustly generates feasible multi-segment screw-motion manipulation sequences. It significantly improves task success rates for intricate operations and overcomes the limitations of traditional “grasp-and-place” paradigms.
To address the trade-off between inefficient trajectory planning and inaccurate grasp success estimation in autonomous multi-fingered grasping, this paper proposes a “plan-then-evaluate” framework. It introduces, for the first time, a vectorized motion planner that enables parallel trajectory generation for multiple candidate grasp poses output by a grasp generator. The evaluation module computes a consistency score over the unified set of planned trajectories, eliminating the time–accuracy compromise inherent in conventional sequential frameworks—where suboptimal grasps are repeatedly re-planned or precision thresholds are relaxed. The framework is agnostic to both grasp generators and motion planners, enhancing system flexibility and real-time performance. Experiments demonstrate significantly higher grasp success rates than baseline methods across diverse objects and realistic settings (e.g., varying shelf and table heights), with strong generalization capability.
This work addresses the challenge of autonomous grasp target selection and accurate 6DoF pose estimation for robots operating in stacked, structured environments (e.g., bricklaying, warehouse stacking), where objects are subject to multi-layer occlusion. We propose the first unified framework jointly optimizing object selection and pose estimation, incorporating a hierarchical selection policy that prioritizes unoccluded top-layer objects. To support systematic evaluation, we introduce the first dedicated benchmark dataset for stacked scenes and define a composite metric integrating selection rationality and pose accuracy. Our method builds upon a tightly coupled camera–IMU perception architecture, synergistically fusing geometric priors with deep learning features to enable robust stack-layer parsing and 6DoF pose regression. Extensive experiments on our custom dataset demonstrate significant improvements over baseline methods. Furthermore, real-world deployment in robotic brick grasping validates the approach’s practicality and reliability under challenging conditions—including variable illumination and partial occlusion.
This work addresses budget-constrained robotic pick-and-place tasks by proposing a low-cost feasibility prediction method that requires neither physics simulation nor real-world interaction supervision. To rank candidate pick-and-place pairs, we design a lightweight dual-output MLP that takes a grasp pose as input and jointly predicts two geometric labels: inverse-kinematics feasibility and path collision risk. We enhance geometric generalization via path-aware pose encoding and a waypoint mesh sweep template, and adopt a “rank-then-plan” strategy for efficient resource allocation. Our approach is the first to achieve cross-domain transfer under purely geometric supervision. On real hardware, it significantly reduces planner invocations, accelerates discovery of successful paths, and maintains—or even improves—task success rates. This establishes an efficient, unsupervised learning paradigm for low-budget robotic manipulation.
This work addresses the common oversight in existing grasping methods of spatially non-uniform mechanical properties on object surfaces, which often leads to damage when contacts occur at fragile regions. The authors propose a novel grasping framework that integrates language instructions, 3D reconstruction, and physical awareness: leveraging SAM3D for language-guided 3D reconstruction, they perform physics-informed geometric analysis to generate local contact force tolerance maps, which are then used to filter and re-rank candidate grasp poses according to task consistency and force-map awareness. During execution, an adaptive impedance controller dynamically modulates finger stiffness based on contact points. This approach is the first to incorporate local mechanical tolerance maps throughout the entire grasping pipeline, reframing dexterous manipulation from a purely geometric problem into a joint optimization under physical constraints. Experiments demonstrate stable selection of high-strength contact regions and maintenance of grasp forces within safe thresholds on paper, plastic, and glass cups.
Robot grasping models trained on limited real-world data struggle to generalize to geometrically diverse novel objects in industrial settings. Method: We introduce GraspNet-1B, the first large-scale, object-centric synthetic grasping dataset, comprising over 100 million 6-DoF grasp poses across tens of thousands of 3D objects and major grippers (e.g., Franka Panda, Robotiq 2F-85), generated via high-fidelity physics simulation to maximize geometric diversity and cross-object generalization. Contribution/Results: Models trained on GraspNet-1B achieve significantly higher grasp success rates on unseen objects in both simulation and real-robot experiments. The dataset establishes a scalable, standardized benchmark for data-driven grasping generalization in industrial applications, enabling robust, practical deployment without requiring extensive real-world annotation.
Sampling-based 3D motion planning for rigid bodies in six-dimensional configuration space suffers from low sampling efficiency in narrow passages and fails to reuse historical planning experience. Method: This paper proposes a reusable path library–based planning framework. Its core innovation is the first introduction of cross-object path transfer: historical paths are retrieved via shape similarity matching, then adapted through configuration-space path deformation and path-aligned directed sampling to generate high-quality initial trajectories—thereby substantially reducing the search burden on sampling-based planners (e.g., RRT). The method is implemented and open-sourced within the OMPL framework. Contribution/Results: Experiments demonstrate up to 85% reduction in planning time, significantly improved success rates in narrow-passage scenarios, and successful solutions for several complex cases where all baseline methods fail.
This work addresses the misalignment between the confidence scores produced by existing 6-DoF grasp detectors and actual grasp quality, which often leads to high-quality candidates being underestimated. To resolve this, the paper introduces a novel grasp re-ranking module decoupled from the detector, enabling refinement without modifying the frozen detector. By integrating candidate attributes, shell-wise hierarchical local geometry, and object-level contextual information through a conditional representation and a Transformer-based multimodal fusion strategy, the method predicts more accurate grasp quality scores for improved ranking. Evaluated on GraspNet-1Billion, the approach boosts the average AP of three frozen detectors by up to 13.60 points. Real-world robotic experiments further demonstrate its robust performance in cluttered scenes.
To address the trade-off between computational efficiency and modeling expressiveness in dynamic modeling of high-dimensional deformable objects, this paper proposes a task-driven, spatially adaptive method for automatic dynamic model generation. The method introduces diffusion models to predict local modeling resolution at critical regions from point-cloud-based planning queries—a novel application in this domain. It further designs a two-stage joint optimization framework that unifies predicted dynamics priors with closed-loop control performance feedback, enabling co-optimization of spatial resolution distribution and dynamical parameters. Integrating point-cloud perception, differentiable physics simulation, and closed-loop simulation data collection, the approach achieves a 2× speedup over full-resolution modeling on tree-like object manipulation tasks, with less than 3% degradation in task success rate. This substantially enhances both planning efficiency and practical deployability.