Score
Designs and implements algorithms and pipelines that estimate inter-frame motion fields and apply spatial transformations to align frames across time and stabilize landmarks and regions. This includes methods for motion estimation, registration, warping, and temporal smoothing to compensate for head and facial motion.
This work addresses key bottlenecks in motion sequence temporal alignment—namely, reliance on paired data, cross-domain mapping, and supervised training. We propose a zero-shot, multimodal motion matching framework that performs alignment solely via metric distances among motion patches within a single domain, leveraging local distance computation and optimal transport—without any cross-domain modeling or labeled supervision. The framework supports diverse control inputs, including sketches, semantic labels, audio, or reference motions. Critically, it enables robust and efficient motion retargeting without requiring paired samples or supervised training. Experiments demonstrate substantial reductions in both data acquisition and computational costs. Comprehensive evaluations across multiple control tasks confirm strong generalization capability and practical deployability.
为解决帧插值中的运动模糊问题,提出ABC-Inter算法,通过贝塞尔控制点和精确流估计模块提高光流估计准确性。
Medical image registration across heterogeneous imaging devices suffers from fragmented and nontraceable transformation information, impeding clinical diagnosis and collaborative workflows. To address this, we propose a tree-structured documentation framework for multimodal image registration, unifying coordinate transformations—spanning diverse devices and modalities—within a patient-specific reference frame. We introduce the .dpw (Digital Patient Workspace) proprietary file format, enabling hierarchical storage, reversible provenance tracking, and cross-platform reproducibility of transformation chains. Furthermore, we develop dpVision, a software tool supporting interactive visualization, validation, and management of registration workflows. Evaluated in orthodontic analysis, our approach significantly enhances interpretability, reproducibility, and clinical audit efficiency of complex registration pipelines. The method provides a standardized, interoperable foundation for multicenter medical imaging collaboration.
Existing diffusion-based human image animation methods suffer significant degradation in generation quality when scale or rotational misalignment exists between the reference image and target pose—severely limiting practical applicability. To address this, we propose Test-time Procrustes Calibration (TPC), the first geometric calibration technique integrated into diffusion-based animation frameworks. TPC performs training-free, plug-and-play geometric realignment of the reference image at test time via Singular Value Decomposition (SVD)-based Procrustes analysis, and adaptively adjusts the diffusion model’s conditional inputs accordingly. The method is model-agnostic, incurs zero training overhead, and introduces no additional inference latency. Extensive evaluations across multiple benchmarks demonstrate that TPC reduces FID by 23.6% and improves keypoint consistency by 31.4%, effectively mitigating the pervasive composition misalignment problem in real-world scenarios.
4D medical image interpolation faces a fundamental trade-off between temporal resolution and reconstruction fidelity. To address this, we propose the first continuous spatiotemporal motion modeling framework inspired by fluid dynamics principles, jointly leveraging Eulerian and Lagrangian descriptions. Our method employs implicit neural representations to ensure both spatial and temporal continuity, enabling training-free, patient-specific optimization. Crucially, it abandons conventional discrete deformation fields in favor of parameter-free forward deformation modeling, thereby substantially improving motion representation accuracy and generalizability. Evaluated on multi-center 4D CT and MRI datasets, our approach achieves average improvements of +3.2 dB in PSNR and +0.04 in SSIM over prior state-of-the-art methods, while operating at 2.1× faster inference speed. Moreover, it requires no large-scale annotated data, making it highly practical for clinical deployment.
This study addresses the inherent trade-off between semantic responsiveness and structural fidelity in text-driven 3D motion editing by proposing the CIME framework. The method introduces a novel spatiotemporal collaborative decoupling mechanism that disentangles variations and invariances into pose and rhythm dimensions. By integrating fully supervised positive-negative learning with Riemannian Non-uniform Integral Manifold Mapping (RNIMM), CIME achieves a refined balance between semantic alignment and physical rhythmic consistency. Experimental results on benchmarks such as MotionFix demonstrate that CIME attains state-of-the-art performance, significantly enhancing both editing alignment and structural preservation. The source code and pretrained models have been made publicly available to facilitate future research.
This work addresses the challenge of correspondence ambiguity in video frame interpolation caused by large-scale nonlinear motion and complex occlusions. To this end, the authors propose a dynamic motion trajectory–based feature scanning mechanism that constructs feature sequences along nonlinear paths guided by optical flow. The method introduces a learnable residual velocity update and a velocity-aware state space model (SSM) to enable adaptive dense sampling and feature aggregation in fast-moving regions. Integrated with an end-to-end jointly optimized intermediate flow estimation and occlusion-aware refinement module, the proposed approach achieves state-of-the-art performance on standard benchmarks, particularly excelling in scenarios involving large displacements and intricate dynamic content.
This work addresses the challenges of unnatural transitions and poor preservation of temporal motion characteristics in motion clip stitching by proposing a learning-free, parameter-free optimization method based on Rodrigues vectors. By representing joint rotations as continuous Rodrigues vectors and formulating the stitching process as a Laplacian smoothing problem in the time domain, the approach effectively enforces rotational continuity and numerical stability—leveraging the observation that rotation axis flips are rare in real human motion. The method supports both intra-class replacement and cross-category motion stitching, producing visually coherent and temporally faithful transitions even between highly dissimilar motions, while enabling efficient interactive editing.
This study addresses platform motion sensitivity and low estimation accuracy in control-point-free long-range visual deformation monitoring by proposing a control-adaptive differential framework. Without requiring nonlinear optimization or initial pose priors, the method employs a decoupled rotation-translation recovery strategy that renders rotation estimation immune to control field contamination while precisely eliminating translational extrinsic errors. Experimental results demonstrate state-of-the-art performance under control-point-free conditions, achieving a rotation RMSE of 2.97 arcseconds with only 0.46 ms computation time, a single-point translation RMSE of 1.19 mm, and a bridge displacement RMSE of 0.85 mm. These findings confirm the framework’s capability for high-precision, efficient relative motion estimation in practical engineering applications.
Existing video editing methods struggle to simultaneously and precisely control both the motion and spatial positioning of subjects, often resulting in distortions or unpredictable outcomes. To address this limitation, this work proposes TeleMorpher—the first one-shot framework enabling joint motion and position editing—by decoupling foreground and background, introducing a training-free pose deformation mechanism, and guiding diffusion model inference with motion priors. Leveraging pretrained segmentation and inpainting models, TeleMorpher achieves high-fidelity edits in a single forward pass. For more reliable evaluation, two novel LPIPS-based metrics are introduced to separately assess background consistency and motion fidelity. Extensive quantitative and subjective evaluations on real-world videos and the TaiChi dataset demonstrate that TeleMorpher significantly outperforms existing approaches, confirming its superior performance and robustness.