Score
Design and implement mechanisms that encode source visual inputs into compact, manipulable tokens and integrate those tokens as conditioning signals for generative models. This includes building instance-level token encoders and pipelines that concatenate visual tokens with text embeddings to provide instance-dependent conditioning for diffusion-based synthesis and enable inference-time token manipulation to produce target-style or labeled images.
Conditional image synthesis with diffusion models suffers from a lack of systematic understanding due to architectural complexity, task heterogeneity, and diverse conditional mechanisms. Method: This paper introduces the first unified taxonomy for conditional diffusion modeling, categorizing approaches by *where* conditioning is injected—either into the denoising network architecture or the sampling process—and formalizes three paradigmatic stages: training, reuse, and specialization. It further classifies six mainstream sampling-time conditioning strategies. Contributions: Based on a structured analysis of over 100 works, the paper establishes a comprehensive knowledge framework and open-sources an authoritative resource repository (GitHub Awesome-Conditional-Diffusion-Models). It identifies persistent bottlenecks—including limited generalization, inefficient inference, and coarse-grained control—and proposes principled directions toward scalable, modular, and fine-grained conditional modeling.
This work proposes CompTok, a framework designed to achieve fine-grained semantic control and enhance model learnability in image generation. By integrating a conditional diffusion decoder with an InfoGAN-inspired recognition objective, CompTok enforces the effective utilization of all visual tokens. During training, subsets of tokens from different images are swapped, and unlabeled adversarial manifold regularization is introduced to preserve generation fidelity, thereby improving token compositionality and controllability. The approach innovatively combines token swapping with manifold constraints and introduces two novel generator-free metrics to evaluate compositionality and learnability in the token space. CompTok achieves state-of-the-art performance on class-conditional image generation, enables high-level semantic editing—such as cross-image token swapping—and significantly outperforms existing methods on the proposed evaluation metrics.
Conventional image editing and generation rely on large, pre-trained generative models, incurring high computational and data requirements. Method: We propose a training-free paradigm based on a one-dimensional vector-quantized tokenizer (1D VQ-Tokenizer) with only 32 discrete tokens, achieving extreme compression (~1024×) while preserving rich semantic structure in the latent space. Fine-grained editing is performed via token-level heuristic operations (e.g., copy, replace), and end-to-end generation is realized through test-time gradient optimization with plug-and-play losses—reconstruction and CLIP-guided similarity. Contribution/Results: We empirically demonstrate, for the first time, that this 1D latent space supports strong semantic editability and generative capability without model training. It enables zero-shot inpainting and text-guided editing. Experiments show competitive diversity and photorealism compared to supervised methods, while drastically reducing computational cost and data dependency.
Existing reference-guided diffusion models suffer from high computational overhead and low inference efficiency under multi-reference conditions. This work proposes Sparse Context, a method that constructs sparse reference representations by introducing a random token dropping strategy during training and enabling task-aware selection of critical tokens at inference time, thereby decoupling token selection from model training. The approach achieves 2× and 4× inference speedups on single-reference and multi-reference generation tasks, respectively, while preserving strong spatial alignment capabilities and high-quality subject-driven generation performance.
This work addresses the challenge of precise control in conditional discrete generative models when confronted with unseen condition combinations. The authors propose a theory-driven, composable discrete generation framework that integrates parallel token prediction with an absorbing diffusion mechanism and a concept-weighted conditional fusion strategy. This approach enables accurate modeling of an arbitrary number and combination of conditions while unifying mask-based generation within the same paradigm. Leveraging compositional vocabularies derived from VQ-VAE/VQ-GAN, the method achieves a 63.4% average reduction in error rate, a 9.58 improvement in FID, and 2.3–12× faster inference across three datasets. Furthermore, it successfully extends to pretrained text-to-image models, enabling fine-grained controllable generation.
Existing vision generation models commonly employ token merging strategies that neglect semantic importance, leading to information loss, degradation of fine details, and generation artifacts. To address this, we propose a semantic-aware dynamic token merging method that— for the first time—systematically leverages importance scores derived from classifier-free guidance to inform merging decisions, prioritizing retention of high-information tokens. Our approach integrates seamlessly into mainstream diffusion models—including Stable Diffusion, Zero123++, AnimateDiff, and PixArt-α—without architectural modification. Extensive experiments across text-to-image, multi-view, and video generation tasks demonstrate substantial improvements: PSNR increases by 3.2 dB, FID decreases by 18.7%, and both detail fidelity and spatiotemporal coherence are significantly enhanced. Moreover, inference speed is accelerated by up to 2.1×. This work establishes a principled, guidance-driven paradigm for adaptive token compression in diffusion-based generative modeling.
This work addresses the challenge of achieving precise and controllable image generation with diffusion models in the absence of large-scale annotated data. The authors propose a self-conditioning mechanism leveraging pretrained self-supervised representations, which identifies semantic directions in the representation space to guide the diffusion process without requiring labeled conditions. This approach not only enhances unconditional generation quality but also constructs a smooth and disentangled controllable generation space. Experimental results demonstrate that the proposed method achieves superior performance in image generation and editing tasks, excelling in controllability, smoothness, and disentanglement compared to existing alternatives.
This study addresses the limitation of existing image and video diffusion models that uniformly allocate computational resources across the entire frame, struggling to adapt to unevenly distributed scene details. To overcome this, we propose a multi-resolution token layout framework that, for the first time, translates prior knowledge into explicit multi-resolution token arrangements, enabling on-demand compute allocation through dynamic adjustment of token granularity. Methodologically, we design a patch-asymmetric flow parameterization alongside a multi-resolution token embedding mechanism, and introduce semantic- and depth-guided layout strategies to enhance generation efficiency while preserving pretrained priors. Experiments demonstrate that our approach achieves substantial acceleration in both image and video generation tasks, maintaining an excellent trade-off between visual quality and computational efficiency.
This study addresses the insufficient high-level semantic planning in Diffusion Transformers for video generation by proposing BiVidGen, a novel framework that introduces an explicit semantic planning paradigm bridged by discrete visual tokens generated via Multimodal Large Language Models. By integrating an EMA tokenizer, autoregressive modeling, and multi-layer cross-attention mechanisms, the method achieves dual-conditioned rendering with both textual and visual tokens. As a systematic exploration of MLLM-DiT integration, BiVidGen significantly improves semantic alignment and temporal coherence on VBench-Long, outperforming fine-tuned DiT baselines and effectively enhancing semantic controllability in long-form video synthesis.
This study addresses the limitation that semantic representations in diffusion models emerge merely as byproducts of generation, making it difficult to enhance representational capacity without compromising generative quality. To overcome this, we propose a cross-view class-token alignment framework for diffusion Transformers that elevates representation learning to a primary optimization objective. The method introduces dual-timestep independent noise observations with an EMA teacher target, integrated with flow matching, self-supervised patch alignment, and stop-gradient mechanisms. We provide the first demonstration that diffusion model representations can be directly optimized rather than solely serving generation. Empirically, our approach yields substantial improvements in ImageNet linear probing accuracy and increases VOC segmentation mIoU by 3.6%, while maintaining FID scores and enhancing text-to-image synthesis quality, thereby achieving synergistic gains in both generation and representation.
This study investigates how reference images and text prompts jointly influence output generation in contextual image synthesis through a unified attention mechanism. Focusing on the multimodal DiT model FLUX.2, the authors employ three causal intervention techniques—T2I Lens, Attention Knockout, and I2I-to-I2I Patching—to reveal, for the first time, that text tokens, particularly filler tokens, serve as structured conduits for visual reference information. They further discover that pixel-level identity information can bypass textual mediation entirely via image-to-image attention, establishing a dual-path transmission mechanism. Experiments across 2,875 editing tasks demonstrate that attributes such as color and style are conveyed through text tokens, whereas specific instance identity relies on the image pathway, thereby clarifying the division of labor between semantic and identity information in multimodal generative models.