Score
Designs, implements, and analyzes mathematical mappings that reproject points, vectors, and poses between coordinate frames — including change-of-basis operations, rotation and translation matrices, and spherical/cartesian representations. Tasks include deriving and composing transforms from sensor poses, implementing reprojection and coordinate-mapping algorithms (including differentiable variants), and applying those transforms to convert data between frames.
This work addresses the problem of establishing affine transformation relationships between local image patches observed by two calibrated cameras. By integrating multi-view geometry with differential geometric analysis, the authors derive the first closed-form solution for the affine transformation that is valid for arbitrary calibrated camera pairs. The solution explicitly characterizes the analytical dependence of the transformation on the relative pose, image coordinates, and local surface normal. The proposed method is not only concise in formulation and computationally efficient but also provides a rigorous theoretical foundation for applications such as image registration, feature matching, and 3D reconstruction.
This work addresses the challenge of generalizing imitation-learning-based manipulation policies across objects with significant geometric discrepancies—e.g., pouring liquid into unseen containers with novel shapes and poses. We propose the Motion Transfer Frame (MTF) framework, which automatically identifies geometry-agnostic key points and dynamic reference frames grounded in both object geometry and task semantics. MTF integrates geometric-aware keypoint localization, reference-frame binding, and kinematic constraint modeling, and supports closed-loop validation from simulation to real robots. Its core contribution is geometry-invariant trajectory transfer that simultaneously enforces critical pose constraints (e.g., cup upright orientation), collision-free motion, and task success. Experiments demonstrate >92% pouring success across diverse unseen container configurations, substantially improving cross-morphology generalization of manipulation skills.
Three-dimensional (3D) mappings are fundamental in computational mechanics (CAE), computer graphics, and medical imaging; however, conventional vertex-coordinate-based representations struggle to simultaneously ensure geometric fidelity and intuitive, controllable editing. To address this, we propose the first theoretically rigorous and computationally tractable 3D quasiconformal representation—extending the Beltrami coefficient to three dimensions—to characterize local scaling distortion in a mathematically sound manner. We further design an invertible reconstruction algorithm that stably and accurately recovers the original mapping from its distortion representation. Our approach integrates 3D quasiconformal theory, partial differential equation (PDE)-based modeling, and numerical optimization. Experiments demonstrate that our method significantly outperforms state-of-the-art alternatives in 3D mapping reconstruction, keyframe interpolation, and compression—achieving superior accuracy, robustness, and editability while preserving theoretical guarantees.
This work addresses the lack of physical plausibility in single-image 3D reconstruction. We propose the first physics-compatible reconstruction framework that enforces static equilibrium as a hard constraint. Methodologically, we explicitly decouple and jointly optimize material stiffness, external loading forces, and the static equilibrium geometry; deformation responses are modeled via differentiable physics simulation, enabling gradient-based joint optimization of all variables. Our approach breaks from conventional simplifications—such as rigid-body assumptions or neglect of external forces—by embedding real-world physical constraints directly into the single-image reconstruction pipeline. Evaluated on Objaverse, our method yields reconstructions with significantly improved mechanical stability, suitable for downstream dynamic simulation and 3D printing. Physical validation via real-world force testing further confirms the structural robustness of the generated models.
Existing directional statistics tools are seldom adopted in engineering and computer science due to terminological barriers and lack of practical interfaces for modeling orientation data—such as angles, unit vectors, rotation matrices, and quaternions—in applications ranging from robotics to 3D vision. Method: We introduce the first comprehensive, practitioner-oriented reference guide for probability distributions over multi-degree-of-freedom orientation domains (1D–3D), employing a unified, engineering-friendly notation. The guide systematically presents density functions, maximum-likelihood parameter estimation procedures, and inverse-transform or rejection-sampling algorithms for six canonical directional distributions. Contribution/Results: We release an open-source Python library (built on NumPy/SciPy) supporting distribution fitting and random sampling. Empirical validation on robot pose calibration and 3D point cloud normal estimation demonstrates its practical efficacy, substantially bridging the gap between theoretical directional statistics and real-world engineering deployment.
This work proposes a parameter-free local topographic descriptor for the comparison and rigid alignment of three-dimensional structured point patterns. The method decomposes each point pattern into multiple arms and introduces a normalized finite difference operator along each arm to capture the local variation of height components relative to the underlying planar geometry, thereby integrating fine-grained geometric details with global structural information. By combining Wasserstein distance with Procrustes analysis, the approach enables efficient distributional comparison and precise alignment of point clouds. The proposed descriptor preserves salient local topographic features while significantly enhancing the robustness and accuracy of point pattern matching.
This work addresses the limitations of traditional 3D reconstruction methods, which predict point maps in camera-centered coordinates, struggle to incorporate scene structural priors, and suffer from high rotational degrees of freedom across views, leading to inconsistent reconstructions. To overcome these issues, the authors propose predicting point maps in a gravity-aligned upright coordinate system, thereby reducing inter-view rotational ambiguity through a shared vertical axis. They introduce the Gravity Grounded Geometry Transformer (G3T) model and the G3T-Long incremental reconstruction framework, which for the first time integrate gravity-aligned coordinates into point map prediction by combining a Transformer architecture, gravity-aware pose estimation, and a submap stitching strategy. Experiments demonstrate that this approach significantly improves reconstruction accuracy and robustness, outperforming existing methods in incremental 3D reconstruction and validating the effectiveness of gravity-aligned representations.
This work addresses optimization problems defined over products of simplices, such as low-rank learning of discrete multivariate probability distributions and function data registration based on the Square-Root Velocity Function (SRVF) representation. To tackle the inherent constraints, the authors propose a smooth reparameterization that is strictly convex element-wise, transforming the constrained problem into an unconstrained optimization over a Riemannian manifold. The resulting problem is solved via Riemannian gradient descent (RGD). Theoretical analysis shows that this reparameterization maps second-order KKT points on the manifold to weak second-order KKT points of the original problem, ensuring theoretical soundness while enhancing computational efficiency. Experiments demonstrate that RGD significantly outperforms projected gradient descent (PGD), achieving more accurate shape-preserving registration in functional data and efficiently solving probability tensor decomposition tasks.
This work addresses the lack of reproducible and quantifiable evaluation benchmarks in existing digital twin generation methods, which often rely on subjective qualitative comparisons. To this end, the paper proposes a synthetic image generation framework based on high-fidelity 3D models and programmable camera poses, enabling systematic quantitative assessment of reconstruction results under known ground-truth parameters. The approach introduces, for the first time, a programmable virtual environment coupled with a ground-truth parameter reference mechanism, integrating procedural trajectory generation, photorealistic rendering, and feature-point triangulation-based reconstruction. This framework establishes the first benchmark for digital twin evaluation that supports reproducible and objective comparisons, significantly enhancing the consistency and scientific rigor of assessments across different generation strategies.