GroupVideo: Multi-Identity Customized Text-to-Video Generation

πŸ“… 2026-07-23
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This work addresses the challenges of identity confusion, unnatural motion, and β€œcopy-paste” artifacts in existing personalized text-to-video generation methods when handling multi-identity scenarios. To overcome these limitations, the authors propose a video diffusion Transformer-based framework for multi-identity customization, incorporating multimodal identity alignment (visual and semantic), a spatially guided identity localization module, and a semantic-aware controller to enhance identity disentanglement and motion naturalness. The approach further introduces joint multi-face encoding, bounding box constraints, and mask regularization loss, alongside the construction of the first large-scale multi-identity video dataset. Experimental results demonstrate that the proposed method significantly outperforms current state-of-the-art approaches in both identity consistency and motion naturalness, enabling high-fidelity generation of videos featuring multiple coordinated characters.
πŸ“ Abstract
Current identity customized video generation methodologies are predominantly limited to single-identity scenarios, as the lack of explicit identity separation mechanisms often leads to identity confusion in multi-identity settings. Existing multi-identity approaches, which directly extend single-identity frameworks by concatenating face images as input conditions, frequently result in unnatural facial expressions and motions, manifesting as the "copy-paste" phenomenon. To overcome these limitations, we introduce GroupVideo, a novel framework that leverages multiple individual photographs to generate identitycustomized video. Built upon Video Diffusion Transformers, GroupVideo incorporates multimodal identity alignment: visual alignment jointly encodes multiple face images to provide robust identity references, while semantic alignment introduces a semantic perceiver to enhance the naturalness of motions. An ID localization module with spatial guidance is introduced to address identity blending and enhance identity fidelity, along with bounding box constraints and mask regularization loss, to focus on facial regions and improve training efficiency. In response to the shortage of multi-ID video datasets, we have curated a comprehensive high-quality dataset of 20,000 videos, thereby establishing a crucial resource to advance future research in multi-ID video generation. Extensive experiments demonstrate that GroupVideo outperforms existing methods in generating multi-character videos with consistent identities and natural motions.
Problem

Research questions and friction points this paper is trying to address.

multi-identity video generation
identity confusion
text-to-video
facial motion naturalness
identity separation
Innovation

Methods, ideas, or system contributions that make the work stand out.

multi-identity video generation
video diffusion transformer
multimodal identity alignment
ID localization module
customized text-to-video
πŸ”Ž Similar Papers
X
Xinyang Song
School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing 100049, China; New Laboratory of Pattern Recognition and State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences, Beijing 100190, China
L
Libin Wang
Ant Group, Beijing 100081, China
J
Jianxin Sun
Ant Group, Beijing 100081, China
Qi Li
Qi Li
Institute of Automation, Chinese Academy of Sciences
pattern recognitioncomputer vision
D
Dandan Zheng
Ant Group, Beijing 100081, China
J
JingDong Chen
Ant Group, Beijing 100081, China
Zhenan Sun
Zhenan Sun
Institute of Automation, Chinese Academy of Sciences
BiometricsPattern RecognitionComputer Vision