AnyAct: Towards Human Reenactment of Character Motion From Video

📅 Unknown Date
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge of directly generating editable and physically plausible human motion reenactments from monocular videos of non-human characters, rather than reconstructing the original character’s motion. To this end, the authors propose a cross-structure motion transfer framework based on sparse local 2D joint movements, which operates without requiring a 3D model or topological priors of the source character. The method introduces a global–local motion disentanglement mechanism and integrates an enhanced 3D-to-2D projection supervised solely by human motion data, combined with a progressive training strategy. Experiments on a newly curated benchmark of non-human character videos demonstrate that the approach produces high-fidelity human motions while effectively preserving the original dynamic characteristics, thereby validating its effectiveness and generalization capability.
📝 Abstract
We study the problem of directly deriving an initial human reenactment from a monocular video of a non-human character. Our goal is not to reconstruct the source character itself but to reinterpret its motion as a plausible and editable human performance for downstream animation authoring. This task is challenging because existing video-based motion capture methods are largely restricted to human-centric structural spaces, while motion retargeting methods typically require structured 3D source motions and known source topologies. Our key insight is that sparse local articulated motion cues can preserve essential dynamics across large structural differences, providing a stable bridge from character video to human reenactment. Based on this observation, we propose AnyAct, which formulates character-video-driven human reenactment as conditional human motion generation from transferable sparse local 2D articulated motion. To make this practical, we introduce three key designs: human-motion-only supervision via augmented 3D-to-2D projection, progressive 3D-to-2D training to alleviate conditioning ambiguity, and global-local motion decoupling for reliable local motion control. We further construct a benchmark primarily covering diverse non-human character videos. Experiments on the benchmark show that AnyAct produces high-fidelity initial human reenactments that preserve the essential dynamics of the characters in reference videos, and further ablation studies validate the effectiveness of its core designs.
Problem

Research questions and friction points this paper is trying to address.

human reenactment
character motion
monocular video
motion retargeting
non-human character
Innovation

Methods, ideas, or system contributions that make the work stand out.

human reenactment
motion retargeting
sparse articulated motion
monocular video
conditional motion generation