Vera: Identity-Faithful Human Subject-to-Video Generation

📅 2026-07-22
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Existing video generation methods struggle to maintain identity consistency across frames, especially under significant pose variations and multi-person interactions, often resulting in identity confusion, attribute misalignment, or excessive replication of reference images. To address these limitations, this work introduces Vera, a unified framework built upon the DiT architecture, accompanied by a large-scale, million-level identity-aligned image-video dataset. Vera incorporates Identity-Focused Mask Supervision (IFMS) and Reference-Aware Layered Attention (RALA) mechanisms to enhance identity preservation. The proposed approach significantly improves identity fidelity, subject binding accuracy, and motion naturalness in both single- and multi-person scenarios, effectively mitigating identity ambiguity and overfitting to reference inputs.
📝 Abstract
Subject-to-video (S2V) generation has made substantial progress in preserving reference subjects across diverse categories, yet generic subject consistency remains insufficient for human-centric generation. A video may appear globally consistent while identity-critical human details still drift across frames, poses, and interactions. This issue becomes more severe in multi-person scenarios, where incorrect identity-role binding leads to subject confusion, attribute swapping, and excessive copying of reference-specific appearance cues. We propose Vera, a unified human-centric S2V framework for single- and multi-person generation. We first construct a million-pair identity-aligned human image-video dataset through person-level cross-clip retrieval, providing explicit identity correspondence and diverse references. Built on this dataset, Vera introduces two complementary designs. Identity-Focal Masked Supervision (IFMS) strengthens identity-aware learning with spatially focused supervision while reducing interference from irrelevant artifacts. Reference-Aware Layer-wise Attention (RALA) regulates how video tokens interact with reference identity cues in the DiT backbone, preserving stable identity anchors and enhancing layer-aware identity readout. Extensive experiments demonstrate that Vera improves human identity consistency, multi-person subject binding, and motion naturalness, while reducing identity confusion and excessive reference-image copying.
Problem

Research questions and friction points this paper is trying to address.

subject-to-video generation
identity consistency
human-centric generation
multi-person scenarios
identity-role binding
Innovation

Methods, ideas, or system contributions that make the work stand out.

Subject-to-Video Generation
Identity Consistency
Identity-Focal Masked Supervision
Reference-Aware Layer-wise Attention
Human-Centric Video Synthesis
🔎 Similar Papers
No similar papers found.