Level-of-Token Diffusion

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of existing image and video diffusion models that uniformly allocate computational resources across the entire frame, struggling to adapt to unevenly distributed scene details. To overcome this, we propose a multi-resolution token layout framework that, for the first time, translates prior knowledge into explicit multi-resolution token arrangements, enabling on-demand compute allocation through dynamic adjustment of token granularity. Methodologically, we design a patch-asymmetric flow parameterization alongside a multi-resolution token embedding mechanism, and introduce semantic- and depth-guided layout strategies to enhance generation efficiency while preserving pretrained priors. Experiments demonstrate that our approach achieves substantial acceleration in both image and video generation tasks, maintaining an excellent trade-off between visual quality and computational efficiency.
📝 Abstract
Image and video diffusion models allocate equal computation to every region, even when the intended scene calls for varying levels of detail. The spatial distribution of detail can often be anticipated before generation, indicating where computation can be reduced. We introduce Level-of-Token (LoT) Diffusion, a framework that turns this knowledge into an explicit multiresolution token layout (Level-of-Token layout) for adaptive and efficient generation. Tokens represent rectangular patches of varying sizes and shapes, allocating finer tokens where detail is needed and coarser tokens elsewhere. We adapt pretrained diffusion transformers to LoT layouts through a patch-wise asymmetric flow parametrization and embeddings for multiresolution tokens, preserving full-resolution flow prediction at every denoising step while processing only a reduced token sequence. LoT Diffusion enables layout-adaptive generation while preserving pretrained generative priors. We demonstrate LoT with layouts derived from semantic masks, bounding boxes, texture variance, and depth-of-field cues, as well as agentic plans. Across image and video generation, LoT offers favorable quality-efficiency tradeoffs, with significant speedups determined by the layout's token budget. Our project website is at https://georgenakayama.github.io/lotdiffusion/.
Problem

Research questions and friction points this paper is trying to address.

Diffusion Models
Computational Efficiency
Adaptive Generation
Multi-resolution
Image and Video Generation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Level-of-Token Diffusion
Multiresolution Token Layout
Diffusion Transformer
Asymmetric Flow Parametrization
Adaptive Computation
🔎 Similar Papers
No similar papers found.