SemanTok: Predictable Semantic Tokens for Efficient Autoregressive Video Generation

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the insufficient semantic alignment of flexible-length tokenizers and the low efficiency of autoregressive generation in existing video world models by proposing SemanTok. The method incorporates frozen DINO features into the encoder and employs a lightweight head to independently reconstruct semantic features from each retained token prefix, achieving high semantic consistency across all noise levels via a representation alignment loss. Experimental results demonstrate that SemanTok, with only 201M parameters, performs comparably to VideoFlexTok, which is 3.4 times larger, while significantly improving short-prefix prediction efficiency and video generation fidelity.
📝 Abstract
Recent video-based world models pair the scalability of autoregressive (AR) prediction with the visual quality of diffusion models. The choice of scene tokenizer is paramount for the optimal performance of each of these, both in terms of fidelity and semantics. Flexible-length, coarse-to-fine tokenizers yield exactly that: the first coarse tokens carry the clip's global semantics while later tokens further specify details. Existing flexible tokenizers only apply a representation-alignment (REPA) loss on early decoder hidden states, a target the decoder can partly meet from its noised input instead. We introduce SemanTok, a flexible video tokenizer that feeds frozen DINO features into its encoder and adds lightweight heads that reconstruct them from each retained token prefix alone. SemanTok achieves high semantic alignment and video fidelity at every AR model size: a 201M SemanTok AR model matches or beats a VideoFlexTok AR model $3.4\times$ its size, and larger SemanTok AR models further improve fidelity. It keeps semantic alignment on out-of-distribution classes and gives the decoder higher semantic alignment at every noise level, including pure noise. It performs well in both reconstruction and generation, and its short token prefixes are cheaper to predict and give better generation fidelity, with pixel detail deferred to later tokens.
Problem

Research questions and friction points this paper is trying to address.

video tokenizer
autoregressive video generation
semantic alignment
representation learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Semantic Tokenizer
Autoregressive Video Generation
DINO Features
Representation Alignment
Coarse-to-Fine
🔎 Similar Papers
2024-07-10arXiv.orgCitations: 3