Custom Forcing: Training-Free Subject Customization for Autoregressive Video Generation

📅 2026-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge that subject customization in autoregressive video generation typically relies on costly optimization or bidirectional attention, which is incompatible with causal streaming. To overcome this limitation, we propose a training-free causal streaming customization framework. Specifically, reference frames are stored in a persistent KV cache, while Drift-Adaptive Value Amplification (DVA) and Anchor-Contrastive Guidance (ACG) are introduced to effectively suppress identity drift and mitigate interference from general priors within frozen models. Experimental results demonstrate that the proposed method maintains high subject consistency in two-minute-long video generation, significantly outperforming existing baselines. Furthermore, it achieves a 9.5× to 28.5× speedup in per-frame generation, successfully balancing motion quality with real-time efficiency.
📝 Abstract
Autoregressive video models can generate minute-long videos in real time, but they produce generic subjects from text rather than specific subjects from user-provided images. Existing customization methods either require costly per-subject optimization or use pretrained conditioning networks that jointly process all video frames with bidirectional attention. Neither approach is designed for causal streaming. We present Custom Forcing, a training-free method that stores reference-based anchor frames in the persistent KV cache of a frozen autoregressive video model. However, fixed anchors face two limitations: simple conditioning allows identity to drift, and the text prompt continues to favor a generic subject. To address these problems, drift-adaptive value amplification (DVA) scales reference influence with the degree of identity drift, while anchor contrast guidance (ACG) steers generation away from the generic class prior. Over two-minute rollouts, fixed anchors fall from 0.58 to 0.42 in DINO-I, while Custom Forcing keeps it between 0.58 and 0.62 without reducing motion. Custom Forcing also achieves higher subject similarity than bidirectional customization methods and better preserves identity over 30s than causal image-to-video and reference-to-video models, while generating each frame 9.5--28.5 times faster than these long-video baselines.
Problem

Research questions and friction points this paper is trying to address.

autoregressive video generation
subject customization
identity drift
causal streaming
training-free
Innovation

Methods, ideas, or system contributions that make the work stand out.

Training-Free Customization
Autoregressive Video Generation
KV Cache
Drift-Adaptive Value Amplification
Anchor Contrast Guidance