Persistence Forcing: Exploiting Feature Specialization in Pixel-Space Diffusion

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the imbalance between global structure and local detail representations caused by uniform refinement in pixel-space diffusion models. To this end, we propose Persistent Forcing (PerF), a method that introduces heterogeneous refinement budget allocation and feature grouping mechanisms to encourage natural differentiation of latent representations. Specifically, persistent features encode global structures while active features handle high-frequency details, enabling stable global information to guide detail refinement and providing explicit conditional guidance complementary to classifier-free guidance during generation sampling. Evaluated on ImageNet 256×256, PerF-L achieves an FID of 1.91 with only half the parameter count, while PerF-H attains an FID of 1.63. These results significantly outperform baseline models while preserving high generative quality.
📝 Abstract
Pixel-space diffusion Transformers (DiTs) directly operate on high-dimensional visual data, yet their hidden representations typically undergo uniform refinement across depth. Natural images, however, are inherently organized at different levels of granularity. Global structure can often be represented compactly, whereas local textures and fine details require richer representations. Motivated by this, we introduce heterogeneous refinement in pixel-space DiTs, assigning different feature groups distinct refinement budgets across depth. Consequently, an ordered feature specialization emerges: sparsely refined features predominantly encode global visual structure, whereas more frequently refined features increasingly specialize toward localized, high-frequency details. We refer to these two groups as persistent and active features, respectively. Building on this emergent specialization, we introduce Persistence Forcing (PerF), which explicitly exploits this persistent--active feature organization for pixel-space image generation. This enables persistent features to continuously condition actively refined features, allowing stable global information to guide the ongoing refinement of finer visual details. During generative sampling, this interaction further induces a meaningful guidance direction that promotes coherent global structure and naturally complements classifier-free guidance. On ImageNet $256\times256$, PerF-L achieves FID of $1.91$, approaching $1.86$ of JiT-H with only half the parameters, while PerF-H further achieves FID of $1.63$ and $1.76$ on ImageNet $256\times256$ and $512\times512$, respectively.
Problem

Research questions and friction points this paper is trying to address.

pixel-space diffusion
Diffusion Transformers
heterogeneous refinement
feature specialization
image generation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Pixel-space Diffusion Transformers
Heterogeneous Refinement
Feature Specialization
Persistence Forcing
Classifier-Free Guidance