Vorch-IR: Long-Form Unified Multimodal Identity Replacement Video Generation

๐Ÿ“… 2026-08-06
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This work addresses the limitations of existing video identity replacement methods, which are largely confined to single-person scenarios, rely on structural controls such as pose or masks, and lack paired data for multi-person settings, thereby hindering flexible multimodal editing. The authors propose Vorch-IR, a unified framework built upon the LTX2 model, which jointly conditions on driving videos, indexed reference images, and textual instructions to enable identity swapping for single subjects, pairs, and backgrounds within a single modelโ€”without requiring pose or layout alignment. Key innovations include text-guided reference character selection, an automated paired-data construction pipeline, and a non-autoregressive temporal-overlap inference mechanism capable of generating minute-long videos. Experiments demonstrate that the proposed method significantly outperforms existing approaches in identity fidelity, motion preservation, and temporal coherence.
๐Ÿ“ Abstract
Video identity replacement seeks to transfer the identities of one or more subjects while preserving the motion, expressions, and temporal structure of a driving video. Existing methods largely target single-person settings and often require task-specific structural controls, such as masks or pose representations, limiting their flexibility in general multimodal editing systems. Progress on multi-person replacement is further constrained by the scarcity of paired training data. We present Vorch-IR, a unified framework that supports single- and dual-person identity replacement, with optional background replacement, in a single model. Built on LTX2, Vorch-IR jointly conditions on a driving video, indexed reference images, and a textual editing instruction. The reference images need not match the pose, layout, or spatial configuration of the driving video: their roles as subject or background references are specified through the instruction. Dense visual conditions are fused through self-attention, while a vision-language context establishes semantic correspondence through cross-attention. We further develop an automatic data construction pipeline that synthesizes paired supervision for all four editing settings. Experiments using automatic metrics and pairwise human evaluation demonstrate strong identity preservation, motion fidelity, and temporal coherence across diverse scenarios. A temporal overlapping inference strategy additionally extends the short-clip model to minute-long generation without autoregressive continuation.
Problem

Research questions and friction points this paper is trying to address.

video identity replacement
multimodal editing
multi-person replacement
paired training data
temporal coherence
Innovation

Methods, ideas, or system contributions that make the work stand out.

identity replacement
multimodal video generation
vision-language alignment
long-form video synthesis
unified editing framework
๐Ÿ”Ž Similar Papers
No similar papers found.