Score
Design and implement encoders that convert 3D tool geometry given as point clouds into compact, task‑relevant feature representations usable by control policies, planners, or learning models. Analyze and validate those representations for invariances, robustness, and generalization across different tool shapes and kinematic configurations to support adaptation to changing end‑effectors.
This study addresses the challenge of coupling discrete operation planning with continuous toolpath generation in CNC machining of B-rep models by proposing the CNCGEN framework. This framework introduces a novel "persistent manufacturing object" modeling mechanism that, combined with a learned agent verifier providing material removal feedback, dynamically correlates local predictions with geometric evolution to enable stepwise generation and state updating of operations and toolpaths. The approach integrates deep learning for three-axis machining, B-rep representations, and parametric toolpath algorithms, supported by a synthetically generated dataset incorporating geometric verification. Experimental results demonstrate that, compared to baseline methods, the proposed framework significantly improves workpiece geometric accuracy while effectively mitigating residual material and overcutting phenomena.
To address the high cost of acquiring 3D point cloud data—which limits the scalability of robotic manipulation learning—this paper proposes a vision-based manipulation framework that operates without real 3D inputs. Our core innovation is the 3DStructureFormer module, which transforms monocular RGB images into pseudo-point clouds endowed with explicit geometric structure and employs a dedicated encoder to preserve spatial relationships. We further fuse 2D visual features with these pseudo-3D representations to enable end-to-end policy learning. Evaluated across diverse manipulation tasks—including grasping, pushing/pulling, and insertion—our method achieves performance on par with real-point-cloud baselines, while drastically reducing data acquisition and annotation overhead. Experiments demonstrate that the pseudo-3D representation effectively supports spatial reasoning and generalization. This work establishes a new paradigm for lightweight, deployable vision-manipulation co-learning.
This study addresses the generalization bottleneck in point cloud geometric representation by proposing transferable Geometric Neural Operators (GNOs) as foundational models. Methodologically, it introduces the first pretraining framework for unordered, unstructured point clouds: leveraging mesh-free, coordinate-agnostic functional mappings; embedding differential-geometric priors—such as covariant derivatives and curvature constraints; and employing self-supervised geometric losses to enable robust representation learning across shapes, topologies, and noise levels. Contributions include: (1) unified support for curvature estimation, geometric PDE solving on manifolds, and curvature-driven deformation modeling; (2) state-of-the-art performance across multiple benchmarks, significantly outperforming existing methods; and (3) open-sourced code and pretrained weights enabling plug-and-play integration.
This work addresses the challenge of unstable orientation cues in dexterous hand reorientation tasks due to the unordered nature, inconsistent sampling, and occlusion inherent in point cloud observations. To overcome this, the authors propose a rotation-aware point cloud embedding whose Euclidean distances in the latent space align with the SO(3) geodesic error of object orientations, thereby providing a smooth and geometrically consistent control signal for reinforcement learning policies. Notably, this approach is the first to intrinsically embed rotational geometry at the representation level, eliminating the need for external modules such as pose estimators, optical flow, or teacher supervision. Experiments demonstrate that the method achieves performance on par with baselines relying on privileged state information or distillation, while avoiding their fragile dependence on structured pose estimates or dense optical flow during testing.
This work addresses the challenges of weak spatial understanding and poor cross-view generalization in robotic systems operating under single-view constraints. To this end, the authors propose a unified representation–policy learning framework that introduces a novel single-view 3D pretraining paradigm, integrating point cloud reconstruction with feedforward Gaussian splatting. Through a multi-step knowledge distillation process, geometric-aware representations are effectively transferred to downstream manipulation policies. Evaluated on 12 RLBench tasks, the method achieves an average success rate surpassing the state of the art by 12.7%. Moreover, under large viewpoint shifts, it exhibits only a 29.7% drop in zero-shot success rate—significantly outperforming the state-of-the-art drop of 51.5%—demonstrating its strong viewpoint generalization capability.
This paper presents a systematic review of learning-based 3D representations for tasks such as 3D reconstruction, novel view synthesis, and rendering, spanning from traditional explicit formulations—including meshes, point clouds, and voxels—to emerging implicit neural fields and primitive-based splatting techniques like 3D Gaussian Splatting. Emphasizing the evolutionary trajectory of 3D representations themselves, the work highlights the paradigm shift from explicit to implicit modeling, clarifying the mathematical formulations, strengths, limitations, and suitable application scenarios of each approach. Distinct from prior task-centric surveys, this study centers on representation as the core organizing principle, offering fresh insights for 3D/4D content generation and identifying key challenges and future research directions to serve as a theoretical reference for the computer graphics and vision communities.
This work addresses the challenge of globally characterizing the geometric structure of solution manifolds in redundant robotic tasks, which exhibit non-uniqueness and form continuous manifolds in configuration space. Existing approaches struggle to capture these structures comprehensively. The paper proposes a representation-centric implicit modeling paradigm that constructs a scalar field over the configuration space, whose zero-level set precisely coincides with the task-induced solution manifold. By integrating Jacobian-guided neighborhood sampling with implicit neural representations, the method learns a signed distance field of the solution manifold, enabling globally consistent and continuous modeling under arbitrary task mappings—a capability demonstrated for the first time. Experiments on a planar three-link robot and a seven-degree-of-freedom Franka manipulator validate the approach’s ability to accurately reconstruct solution manifolds and generalize across varying task parameters.
This work addresses the limitations of vision-based robotic manipulation strategies that rely solely on RGB images, which suffer from depth ambiguity and perspective scaling issues, as well as existing point cloud methods that, despite leveraging geometric priors, exhibit poor generalization. The authors propose a novel approach that maps point clouds into a high-dimensional Fourier space for encoding, enabling imitation learning policies to directly perceive high-frequency geometric details and effectively mitigate the neural network’s inherent bias toward low-frequency representations. Evaluated on both simulated and real-world robotic benchmarks—including RoboCasa and ManiSkill3—the method demonstrates significantly improved policy performance, consistently robust results, strong hyperparameter insensitivity, and compatibility with diverse encoding architectures.
Existing video- or partial point cloud–based dynamics models suffer from geometric inconsistency, occlusion sensitivity, and error accumulation over long-horizon predictions, limiting their reliability for planning. This work proposes a task-agnostic, purely 3D world model that tightly integrates point cloud completion with action-conditioned dynamics modeling: it first completes observed partial point clouds into geometrically coherent full scenes and then learns action-driven dynamic evolution directly in the completed 3D space. The approach achieves geometrically consistent and robust long-horizon prediction, enabling stable trajectory forecasting over 100–300+ steps across diverse robotic platforms and tabletop manipulation tasks. It further demonstrates successful sim-to-real transfer, facilitates efficient model-predictive control, and exhibits rapid adaptability to novel tasks.
This study addresses the limitation of behavioral cloning in generalizing to unseen object instances due to overfitting specific geometric appearances. To overcome this, we propose KeyGen, a framework that extracts canonical semantic keypoints from point clouds via unsupervised learning to construct object-centric structured representations. These representations are deeply integrated with visuomotor diffusion policies to ensure cross-instance geometric correspondence consistency and category-level generalization. Additionally, a planning-driven data generation pipeline is designed to establish a simulation benchmark. Experimental results demonstrate that KeyGen significantly outperforms existing methods under pose variations, scale changes, and real-world scenarios. Furthermore, it exhibits favorable scaling behavior with increasing demonstration data, enabling robust robotic manipulation.