ID-V2V: Identity-Preserving Video Restylization

📅 2026-07-24
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge of video re-editing, where preserving identity while enabling flexible manipulation of style, lighting, and scene remains difficult. The authors propose a decoupled approach that formulates the task as a joint optimization of identity-preserving video relighting and keyframe-guided controllable video synthesis. Notably, the method requires only a single source video to construct training pairs, eliminating the need for paired data. By integrating complementary control signals—including relit faces, normal maps, edited keyframes, and depth sequences—into a video-to-video generation framework, the approach effectively preserves fine facial details and ensures temporal consistency. Extensive experiments demonstrate significant improvements over existing methods in terms of facial similarity, micro-expression retention, and visual fidelity, enabling high-quality, temporally coherent editing for both single- and multi-person videos in real-world creative applications.
📝 Abstract
In visual storytelling, human performances are central to creative intent and narrative meaning. However, preserving human identity and performance while enabling flexible visual edits remains challenging for generative video models. We formalize this challenge as identity-preserving video restylization, which propagates scene, lighting, and style changes specified by an edited keyframe across a source video, while preserving facial likeness and performance, including expressions, eye gaze, and lip synchronization. A key obstacle is the absence of paired training data, as identity-preserving restylized video pairs are rare in real-world settings. To address this, we propose a decoupling of source-grounded identity preservation and edit-driven video synthesis. Our key insight is that facial appearance and expression should remain invariant, with illumination being the primary permissible variation. We therefore cast identity preservation as a video relighting problem, while modeling visual edit propagation as controlled video synthesis guided by the edited keyframe. Building on this formulation, we introduce ID-V2V, a video-to-video generative framework integrating complementary control signals: relit facial regions and facial normal maps tightly constrain facial likeness and performance, while edited keyframes and depth sequences enable flexible and temporally coherent generation. This design enables constructing training pairs from a single video, eliminating the need for scarce paired data. Extensive experiments demonstrate that ID-V2V significantly outperforms existing methods in preserving facial likeness and fine-grained facial performance, supports both single- and multi-subject scenarios, and delivers high visual quality, highlighting its potential as a human-centric tool for real-world content production. The code is available at: https://github.com/Eyeline-Labs/ID-V2V.
Problem

Research questions and friction points this paper is trying to address.

identity-preserving
video restylization
facial performance
visual editing
human-centric generation
Innovation

Methods, ideas, or system contributions that make the work stand out.

identity-preserving video restylization
video relighting
controlled video synthesis
facial performance preservation
unpaired training
🔎 Similar Papers
No similar papers found.