WeLike2Party! In-Context Motion Transfer for Multi-Human Image Animation

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of identity-action binding failure and low visual fidelity caused by occlusion in multi-person interactive video generation. To this end, we propose a context-conditioned animation framework that eliminates the need for explicit pose extraction. Methodologically, we introduce the first diffusion-based paradigm for direct video-conditioned generation, incorporating reference-asymmetric RoPE and a contrastive learning-based supervision mechanism for robust identity binding. Furthermore, we construct MotionTwin, a large-scale cross-identity synthetic dataset designed to enhance generalization capabilities. Extensive experiments on both the proposed MotionTwin-Bench and real-world videos demonstrate that our method significantly outperforms existing state-of-the-art approaches in terms of subject visual fidelity, identity-action consistency, and overall perceptual quality.
📝 Abstract
Human image animation aims to transfer motion from a driving video to subjects in a reference image. Despite remarkable progress in video generation, achieving high-fidelity animation of multiple interacting subjects remains a challenge. Many existing approaches rely on explicit motion representations such as 2D skeletons or parametric body meshes and struggle to preserve identity-motion binding under inter-person occlusion. To address this limitation, we propose WeLike2Party, a multi-human animation framework built on direct in-context video conditioning without explicit pose or mesh extraction at inference. We further introduce Reference Asymmetric RoPE Conditioning to preserve fine-grained appearance details, and Identity Binding Supervision to associate each reference identity with its intended motion trajectory. To support cross-identity training, we construct MotionTwin, a large-scale synthetic dataset comprising 14.4K cross-identity video pairs with shared subject and camera motions, totaling 84.3 hours of photorealistic video. We additionally present MotionTwin-Bench, a cross-identity benchmark specifically designed to evaluate subject-level visual fidelity and identity-motion binding. Extensive experiments on MotionTwin-Bench and real-world videos demonstrate that WeLike2Party outperforms recent state-of-the-art methods in subject-level visual fidelity, identity-motion binding, and overall perceptual quality, particularly in multi-person interactions with substantial occlusion.
Problem

Research questions and friction points this paper is trying to address.

multi-human image animation
motion transfer
identity-motion binding
inter-person occlusion
video generation
Innovation

Methods, ideas, or system contributions that make the work stand out.

In-Context Motion Transfer
Multi-Human Image Animation
Reference Asymmetric RoPE Conditioning
Identity Binding Supervision
Cross-Identity Dataset
🔎 Similar Papers
No similar papers found.