Score
Designs and trains a model that represents a continuous, shared deformation field mapping points in a canonical reference to deformed observations across instances or time, capable of expressing non‑rigid spatial transformations. Builds this as a neural or parametric deformation field with shared parameters and optional per‑instance or per‑part latent codes to analyze or synthesize temporal dynamics and instance‑specific variation.
Existing implicit deformation fields suffer from spatial incoherence and poor temporal interpolation capability for dense point trajectory modeling; both neural inductive biases and heuristic explicit methods (e.g., linear blend skinning) struggle to balance interpretability and spatiotemporal consistency. This paper proposes a spline-based explicit trajectory representation: deformation is parameterized via controllable control points, while low-rank time-varying spatial encoding enables decoupled spatiotemporal modeling. The formulation supports analytical computation of velocity and acceleration without imposing rigidity or skinning constraints. Our method significantly improves temporal interpolation accuracy under sparse inputs and enhances motion coherence in dynamic scene reconstruction. It achieves performance competitive with state-of-the-art methods across multiple metrics, while offering superior model interpretability and geometric consistency.
This work addresses the challenge of high-fidelity, temporally consistent 3D motion reconstruction of non-rigid objects—such as clothed humans—from sparse, unstructured, and partially occluded observations. We propose the first end-to-end jointly optimized framework integrating implicit neural fields with explicit mesh deformation. Our method introduces face-level, differential-geometry-driven near-isometric constraints to enforce spatiotemporal consistency in deformations. It further incorporates temporal feature fusion and differentiable rendering optimized by monocular depth video, eliminating reliance on predefined parametric human models. Evaluated on human and animal motion reconstruction tasks, our approach significantly outperforms state-of-the-art methods, achieving substantial improvements in geometric detail preservation and temporal smoothness. Notably, it is the first method to enable high-quality non-rigid motion reconstruction solely from monocular depth video.
This work addresses the challenge of reconstructing dynamic 3D deformations from unordered, sparse, and noisy 4D point cloud sequences. We propose CanFields—the first canonical implicit field framework enabling spatiotemporally consistent modeling of non-rigid motion. Methodologically, we introduce a dynamic consolidator that jointly parameterizes a low-frequency velocity field (ensuring topological integrity and motion continuity) and a high-frequency geometric bias (preserving fine-scale detail), optimized in an unsupervised manner guided by geometric priors—requiring neither explicit supervision nor template initialization. Our key contribution is the first unsupervised formulation that simultaneously guarantees topological stability, geometric fidelity, and temporal coherence. Evaluated on diverse real-world scanned sequences, CanFields reduces reconstruction error by 32% and improves detail preservation by 41% over state-of-the-art methods, while demonstrating superior robustness to occluded regions, sparse frames, and sensor noise.
Existing conditional neural fields (CNFs) suffer from limited performance on fine-grained geometric reasoning tasks—such as classification, segmentation, and reconstruction—due to the lack of explicit modeling of local geometry (e.g., locality, orientation) in their latent spaces. To address this, we propose Equivariant Neural Fields (ENFs), the first CNF framework incorporating *implicit geometric equivariance*. ENFs achieve explicit geometric alignment and equivariant mapping between latent space and continuous signals via geometry-aware cross-attention, coupling neural field decoding with point-cloud–based geometric latent variables. These variables exhibit interpretable rotation/translation covariance, enabling geometric reasoning and local weight sharing. Our method encompasses geometric latent variable modeling, equivariant cross-attention, neural field conditioning, and joint point-cloud–field optimization, efficiently implemented in JAX. Experiments demonstrate that ENFs consistently outperform geometry-agnostic baselines across classification, segmentation, prediction, reconstruction, and generation tasks, significantly improving geometric fidelity of latent representations and generalization efficiency.
This work addresses the geometric quantification and comparison of deformable shapes in images. We propose the Geodesic Deformation Network (GDN), the first method to directly learn geodesic flows—i.e., optimal deformation paths—as end-to-end differentiable mappings from raw images. Methodologically, GDN models geodesics as learnable mapping functions, jointly optimized via a novel geodesic loss; it integrates neural operator architectures, integral operators, smooth activation functions, and explicit constraints from the geodesic differential equation to implicitly parameterize the deformation manifold. Unlike conventional approaches that only estimate initial velocity fields, GDN explicitly learns the entire geodesic path, yielding superior regularity and generalization. Evaluated on 2D synthetic data and 3D real brain MRI, GDN achieves higher deformation alignment accuracy and produces geometrically interpretable, physically meaningful deformation trajectories.
Existing 4D dynamic shape generation methods often suffer from poor temporal consistency and low rendering efficiency due to the tight coupling between shape and motion representations. This work proposes a decoupled 4D representation framework that separates shape and motion into distinct latent spaces, integrating a conditional neural signed distance field with a part-based neural deformation model. By predicting per-part skinning weights and rigid transformations, the method enables structure-aware, efficient modeling of dynamic shapes. It consistently outperforms current state-of-the-art approaches across unconditional generation, conditional generation, and motion retargeting tasks, achieving substantial improvements in generation quality, temporal coherence, and rendering speed.
Existing approaches to generating realistic 4D dynamics—i.e., temporal deformations of objects under varying physical conditions—are constrained by predefined physics-based models, limiting their generalization and scalability. This work proposes Neural Object Kinematics (NeuROK), which reframes object-centric 4D dynamics as a data-driven dynamical system in a low-dimensional latent space, thereby eliminating reliance on task-specific physical priors. Built upon a Transformer encoder–decoder architecture, NeuROK learns a latent representation of object states and their mapping to deformed shapes, enabling efficient and generalizable dynamic generation across a large-scale, cross-category 4D dataset. Experiments demonstrate that NeuROK substantially outperforms current methods and significantly streamlines the pipeline for high-fidelity 4D dynamic synthesis.
This work addresses the challenge of accurately and efficiently modeling continuous deformations of deformable objects, a task where existing methods often sacrifice either precision or real-time performance due to reliance on discrete time steps or high computational costs. The authors propose a novel 4D continuous deformation modeling framework based on Neural Ordinary Differential Equations (Neural ODEs), which maps 3D point clouds and physical conditions into a unified latent space to enable efficient and temporally continuous dynamics simulation. This approach represents the first extension of Neural ODEs to 4D motion modeling of deformable objects, demonstrating strong generalization to unseen shapes and physical parameters, as well as superior interpolation and extrapolation capabilities. Experiments show significant improvements in prediction accuracy under unseen physical configurations, successful transfer to real-world 3D capture data, and the release of code and datasets to support further research.
This study addresses the challenge of accurately characterizing the local second-order statistics of non-stationary random fields induced by spatial deformations. To this end, it proposes a tangent-space covariance model based on local linearization of the deformation mapping and derives, for the first time, a closed-form expression for its local spectral representation. By integrating Gaussian random field simulation with truncated singular value decomposition, the method enables efficient and accurate generation of complex deformation fields. When applied to ACDC cardiac MRI data, the approach successfully uncovers directional and anisotropic differences in myocardial deformation across diagnostic groups, significantly outperforming conventional metrics that rely solely on expansion or compression.
This study addresses the challenge of manual parameter tuning in motion modeling for deformable linear objects by proposing the first automated "video-to-model" framework. The approach integrates structured numerical models with deep learning: a perception module tracks suture trajectories, while a spatiotemporal convolutional neural network automatically estimates parameters for the CBF-CLF-QP control model, enabling end-to-end construction of dynamic models directly from video data. Experimental results demonstrate that the framework achieves high-fidelity reconstruction of thread behavior under unseen configurations, exhibiting low tracking errors and accurately reproducing expected motions. By significantly reducing manual intervention, this work effectively enhances both the efficiency and generalization capability of physics-based modeling.