TeleMorpher: Toward Robust Simultaneous Motion-Location Editing

📅 2026-06-17
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Existing video editing methods struggle to simultaneously and precisely control both the motion and spatial positioning of subjects, often resulting in distortions or unpredictable outcomes. To address this limitation, this work proposes TeleMorpher—the first one-shot framework enabling joint motion and position editing—by decoupling foreground and background, introducing a training-free pose deformation mechanism, and guiding diffusion model inference with motion priors. Leveraging pretrained segmentation and inpainting models, TeleMorpher achieves high-fidelity edits in a single forward pass. For more reliable evaluation, two novel LPIPS-based metrics are introduced to separately assess background consistency and motion fidelity. Extensive quantitative and subjective evaluations on real-world videos and the TaiChi dataset demonstrate that TeleMorpher significantly outperforms existing approaches, confirming its superior performance and robustness.
📝 Abstract
Diffusion models have achieved remarkable success in image and video generation and editing. While recent studies have extended these efforts toward motion editing, simultaneously transforming both motion and location-despite its practical importance-remains largely unexplored. To better understand robust motion-location editing, we first analyze the fundamental factors that degrade its quality. Based on this analysis, we propose TeleMorpher, one of the first one-shot frameworks to the best of our knowledge, for simultaneous motion-location editing. Our approach leverages motion priors, a target motion-centric video generated from an off-the-shelf model as motion-editing guidance, and the ground truth motion to enable more controllable and precise motion-location editing. Via this, our framework works as follows: (1) we first disentangle the protagonist and the background via pre-trained segmentation and inpainting models. (2) Then, we introduce a training-free pose warping that edits the protagonist's motion with the motion prior as the guidance. (3) The result of warped motion video is directly injected into a baseline motion editor during inference, mitigating the difference between source and target motions while preserving the appearance of the source video. (4) To enhance the reliability of quantitative evaluations, we propose two new LPIPS-based metrics that measure the background consistency before and after the motion editing and the fidelity of motion editing performance via measuring the difference between the extracted protagonist's skeletons from source and target videos. Experiments with in-the-wild videos and the TaiChi dataset demonstrate that TeleMorpher achieves superior performance across both quantitative and qualitative measurements (real-human evaluation), underscoring its effectiveness.
Problem

Research questions and friction points this paper is trying to address.

motion editing
location editing
simultaneous editing
video generation
diffusion models
Innovation

Methods, ideas, or system contributions that make the work stand out.

motion-location editing
diffusion models
pose warping
training-free
video editing
🔎 Similar Papers
No similar papers found.