visual token conditioning

Design and implement mechanisms that encode source visual inputs into compact, manipulable tokens and integrate those tokens as conditioning signals for generative models. This includes building instance-level token encoders and pipelines that concatenate visual tokens with text embeddings to provide instance-dependent conditioning for diffusion-based synthesis and enable inference-time token manipulation to produce target-style or labeled images.

visualtokenconditioning

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.33
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This work proposes CompTok, a framework designed to achieve fine-grained semantic control and enhance model learnability in image generation. By integrating a conditional diffusion decoder with an InfoGAN-inspired recognition objective, CompTok enforces the effective utilization of all visual tokens. During training, subsets of tokens from different images are swapped, and unlabeled adversarial manifold regularization is introduced to preserve generation fidelity, thereby improving token compositionality and controllability. The approach innovatively combines token swapping with manifold constraints and introduces two novel generator-free metrics to evaluate compositionality and learnability in the token space. CompTok achieves state-of-the-art performance on class-conditional image generation, enables high-level semantic editing—such as cross-image token swapping—and significantly outperforms existing methods on the proposed evaluation metrics.

compositionalityimage generationlearnability

Highly Compressed Tokenizer Can Generate Without Training

Jun 09, 2025
LL
L. L. Beyer
🏛️ MIT | Meta

Conventional image editing and generation rely on large, pre-trained generative models, incurring high computational and data requirements. Method: We propose a training-free paradigm based on a one-dimensional vector-quantized tokenizer (1D VQ-Tokenizer) with only 32 discrete tokens, achieving extreme compression (~1024×) while preserving rich semantic structure in the latent space. Fine-grained editing is performed via token-level heuristic operations (e.g., copy, replace), and end-to-end generation is realized through test-time gradient optimization with plug-and-play losses—reconstruction and CLIP-guided similarity. Contribution/Results: We empirically demonstrate, for the first time, that this 1D latent space supports strong semantic editability and generative capability without model training. It enables zero-shot inpainting and text-guided editing. Experiments show competitive diversity and photorealism compared to supervised methods, while drastically reducing computational cost and data dependency.

Enables image editing via heuristic token manipulationExplores 1D image tokenizers for high compressionGenerates images without training using optimization

Existing reference-guided diffusion models suffer from high computational overhead and low inference efficiency under multi-reference conditions. This work proposes Sparse Context, a method that constructs sparse reference representations by introducing a random token dropping strategy during training and enabling task-aware selection of critical tokens at inference time, thereby decoupling token selection from model training. The approach achieves 2× and 4× inference speedups on single-reference and multi-reference generation tasks, respectively, while preserving strong spatial alignment capabilities and high-quality subject-driven generation performance.

computational efficiencydiffusion modelsinference speed

This work addresses the challenge of precise control in conditional discrete generative models when confronted with unseen condition combinations. The authors propose a theory-driven, composable discrete generation framework that integrates parallel token prediction with an absorbing diffusion mechanism and a concept-weighted conditional fusion strategy. This approach enables accurate modeling of an arbitrary number and combination of conditions while unifying mask-based generation within the same paradigm. Leveraging compositional vocabularies derived from VQ-VAE/VQ-GAN, the method achieves a 63.4% average reduction in error rate, a 9.58 improvement in FID, and 2.3–12× faster inference across three datasets. Furthermore, it successfully extends to pretrained text-to-image models, enabling fine-grained controllable generation.

composed conditionsconditional compositioncontrollable image generation

Importance-Based Token Merging for Efficient Image and Video Generation

Nov 23, 2024
HW
Haoyu Wu
🏛️ Stony Brook University | EPFL

Existing vision generation models commonly employ token merging strategies that neglect semantic importance, leading to information loss, degradation of fine details, and generation artifacts. To address this, we propose a semantic-aware dynamic token merging method that— for the first time—systematically leverages importance scores derived from classifier-free guidance to inform merging decisions, prioritizing retention of high-information tokens. Our approach integrates seamlessly into mainstream diffusion models—including Stable Diffusion, Zero123++, AnimateDiff, and PixArt-α—without architectural modification. Extensive experiments across text-to-image, multi-view, and video generation tasks demonstrate substantial improvements: PSNR increases by 3.2 dB, FID decreases by 18.7%, and both detail fidelity and spatiotemporal coherence are significantly enhanced. Moreover, inference speed is accelerated by up to 2.1×. This work establishes a principled, guidance-driven paradigm for adaptive token compression in diffusion-based generative modeling.

Allocating computational resources to critical tokens using importance scoresImproving token merging for efficient image and video generationPreserving high-information tokens to enhance sample quality

Latest Papers

What's happening recently
View more

This work addresses the challenge of achieving precise and controllable image generation with diffusion models in the absence of large-scale annotated data. The authors propose a self-conditioning mechanism leveraging pretrained self-supervised representations, which identifies semantic directions in the representation space to guide the diffusion process without requiring labeled conditions. This approach not only enhances unconditional generation quality but also constructs a smooth and disentangled controllable generation space. Experimental results demonstrate that the proposed method achieves superior performance in image generation and editing tasks, excelling in controllability, smoothness, and disentanglement compared to existing alternatives.

controllable image generationdiffusion modelsgeneration control

This study addresses the limitation of existing image and video diffusion models that uniformly allocate computational resources across the entire frame, struggling to adapt to unevenly distributed scene details. To overcome this, we propose a multi-resolution token layout framework that, for the first time, translates prior knowledge into explicit multi-resolution token arrangements, enabling on-demand compute allocation through dynamic adjustment of token granularity. Methodologically, we design a patch-asymmetric flow parameterization alongside a multi-resolution token embedding mechanism, and introduce semantic- and depth-guided layout strategies to enhance generation efficiency while preserving pretrained priors. Experiments demonstrate that our approach achieves substantial acceleration in both image and video generation tasks, maintaining an excellent trade-off between visual quality and computational efficiency.

Adaptive GenerationComputational EfficiencyDiffusion Models

This study addresses the insufficient high-level semantic planning in Diffusion Transformers for video generation by proposing BiVidGen, a novel framework that introduces an explicit semantic planning paradigm bridged by discrete visual tokens generated via Multimodal Large Language Models. By integrating an EMA tokenizer, autoregressive modeling, and multi-layer cross-attention mechanisms, the method achieves dual-conditioned rendering with both textual and visual tokens. As a systematic exploration of MLLM-DiT integration, BiVidGen significantly improves semantic alignment and temporal coherence on VBench-Long, outperforming fine-tuned DiT baselines and effectively enhancing semantic controllability in long-form video synthesis.

Diffusion TransformersMLLM-DiT FusionMultimodal Large Language Models

This study addresses the limitation that semantic representations in diffusion models emerge merely as byproducts of generation, making it difficult to enhance representational capacity without compromising generative quality. To overcome this, we propose a cross-view class-token alignment framework for diffusion Transformers that elevates representation learning to a primary optimization objective. The method introduces dual-timestep independent noise observations with an EMA teacher target, integrated with flow matching, self-supervised patch alignment, and stop-gradient mechanisms. We provide the first demonstration that diffusion model representations can be directly optimized rather than solely serving generation. Empirically, our approach yields substantial improvements in ImageNet linear probing accuracy and increases VOC segmentation mIoU by 3.6%, while maintaining FID scores and enhancing text-to-image synthesis quality, thereby achieving synergistic gains in both generation and representation.

Diffusion ModelsGeneration QualityRepresentation Learning

This study investigates how reference images and text prompts jointly influence output generation in contextual image synthesis through a unified attention mechanism. Focusing on the multimodal DiT model FLUX.2, the authors employ three causal intervention techniques—T2I Lens, Attention Knockout, and I2I-to-I2I Patching—to reveal, for the first time, that text tokens, particularly filler tokens, serve as structured conduits for visual reference information. They further discover that pixel-level identity information can bypass textual mediation entirely via image-to-image attention, establishing a dual-path transmission mechanism. Experiments across 2,875 editing tasks demonstrate that attributes such as color and style are conveyed through text tokens, whereas specific instance identity relies on the image pathway, thereby clarifying the division of labor between semantic and identity information in multimodal generative models.

cross-modal attentionin-context image generationmultimodal DiT

Hot Scholars

JZ

Jiazhao Zhang

Peking University
Embodied AINavigation3D Vision
MY

Mengping Yang

East China University of Science and Technology
Few-shot LearningGenerative Models
YM

Yunze Man

University of Illinois Urbana-Champaign
RoboticsMachine LearningComputer VisionAutonomous Driving
CS

Carmelo Sferrazza

UC Berkeley
RoboticsArtificial IntelligenceHumanoidsTactile Sensing