pixel-space diffusion

Designs, builds, or analyzes generative and denoising diffusion models that operate directly in image/pixel/feature and point‑map representations (including image‑space, pixel‑space, feature‑space, point‑map, and spatio‑temporal/4D variants) rather than in compressed latents. This work covers training, sampling, and evaluation methods that produce or refine raw pixels and point maps in a single stage—bypassing lossy latent compression—to recover sharper geometric and temporal structure and remain robust in ambiguous regions.

pixel-spacediffusion

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.28
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Diffusion Models: A Comprehensive Survey of Methods and Applications

Sep 02, 2022
LY
Ling Yang
🏛️ Peking University | University of California, Los Angeles | Carnegie Mellon University | University of California, Merced

Despite their remarkable performance, diffusion models lack a systematic survey and a unified taxonomic framework. Method: This paper introduces the first comprehensive taxonomy encompassing methodological evolution and cross-domain applications, systematically reviewing over 300 seminal works published between 2015 and 2024. It focuses on three core research directions: efficient sampling, improved likelihood estimation, and modeling of structured data—covering key techniques including denoising score matching, stochastic differential equation (SDE) solvers, latent-space distillation, conditional guidance, and cross-modal joint modeling. Contribution/Results: We propose a novel integration paradigm that synergizes diffusion models with other generative paradigms. Furthermore, we publicly release a structured literature repository and a dynamically updated classification system, which has become a standard reference resource in the field.

Categorizing research into efficient sampling and likelihood estimationReviewing interdisciplinary applications across scientific fieldsSurveying diffusion models' methods and applications comprehensively

Must-Read Papers

Most classic and influential ideas
View more

This study addresses the low training efficiency of pixel-space diffusion models by proposing a latent-to-pixel space transfer strategy. Through systematic optimization of weight initialization, data composition, noise scheduling, and decoder architecture, we establish an efficient pixel-space training paradigm. Experiments demonstrate that this approach effectively resolves slow pre-training convergence, achieving performance that matches or surpasses latent-space models while improving end-to-end inference speed by 3.18× to 4.75×. Consequently, this work provides a comprehensive guideline for high-performance and practical pixel-level training in high-resolution image generation, bridging the gap between theoretical efficacy and computational efficiency in generative modeling.

Large-scale trainingPixel-space diffusion modelsText-to-image generation

Towards Spatially Consistent Image Generation: On Incorporating Intrinsic Scene Properties into Diffusion Models

Aug 14, 2025
HL
Hyundo Lee
🏛️ Seoul National University | Columbia University

Current image generation models often produce spatially inconsistent outputs and geometric distortions due to the lack of explicit scene-structure modeling. To address this, we propose a collaborative denoising framework that jointly generates images and their intrinsic attributes—namely depth maps and semantic segmentation masks—thereby implicitly learning geometric and layout constraints through a shared latent space. Our method builds upon a pre-trained latent diffusion model and employs a lightweight autoencoder to fuse multi-source intrinsic attributes as structural priors. A cross-domain information-sharing mechanism enables synchronized denoising across the image and attribute domains, requiring neither 3D supervision nor additional annotations. Experiments demonstrate that our approach preserves text-image alignment and visual fidelity while significantly improving spatial plausibility. It achieves state-of-the-art performance across quantitative metrics—including layout consistency and depth fidelity—outperforming all baseline methods.

Co-generating images and intrinsic scene propertiesEnhancing scene structure without degrading image qualityImproving spatial consistency in image generation models

There and Back Again: On the relation between Noise and Image Inversions in Diffusion Models

Apr 11, 2026
ŁS
Łukasz Staniszewski
🏛️ Warsaw University of Technology | IDEAS NCBR | Institute of Mathematics | Polish Academy of Sciences | University of Warsaw

Diffusion models achieve high-quality generation but suffer from severe deviation of inverted latent representations—e.g., those obtained via DDIM inversion—from the Gaussian prior, leading to ill-posed latent-to-image mapping and poor diversity in interpolation and editing. This work is the first to systematically identify the root cause: amplified noise prediction errors in smooth image regions, which distort latent-space structure. Through noise prediction error visualization, quantitative evaluation of latent-space editability, and statistical analysis of structural patterns, we empirically confirm systematic structural biases in inverted latents. Our analysis provides an interpretable diagnostic framework for latent-space distortion and, theoretically, establishes essential constraints for constructing semantically consistent and highly controllable diffusion latent spaces. These insights lay a principled foundation for designing next-generation controllable generation and editing methods.

Analyzes relation between noise and image inversions in diffusion models.Demonstrates limitations in diversity and manipulation of image inversions.Explains structural patterns in latents and noise prediction errors.

Diffusion Models as Cartoonists: The Curious Case of High Density Regions

Nov 02, 2024
RK
Rafal Karczewski
🏛️ Aalto University | YaiYai Ltd

Diffusion models unexpectedly generate cartoonized or blurry images—nonexistent in training data—within high-density regions of the learned distribution. Method: We propose Mode Tracking Theory to precisely localize modes in the diffusion denoising distribution; design a zero-overhead SDE likelihood tracking method that estimates and optimizes sample likelihood without additional computation; and develop an efficient high-density sampler that targets atypical, high-likelihood samples overlooked by conventional samplers. Results: Experiments demonstrate substantial improvement in sampling likelihood, stable generation of cartoon/blurry high-density images, and faithful reproduction of this phenomenon on purely real-image datasets—revealing an intrinsic, implicit structural bias inherent to diffusion models.

Develop mode-tracking process for denoising distributionIdentify high-density regions in diffusion modelsPropose high-density sampler for higher likelihood images

Diffusion Models in Low-Level Vision: A Survey

Jun 17, 2024
CH
Chunming He
🏛️ Tsinghua University | Sun Yat-sen University | Shanghai Jiao Tong University | Harbin Institute of Technology | Tianyijiaotong Technology Ltd.

To address the lack of systematic surveys and unified modeling frameworks for diffusion models in low-level vision, this paper presents the first comprehensive survey covering over 20 tasks—including image restoration, enhancement, and generation. We propose three general-purpose diffusion modeling paradigms, theoretically unify them with GANs and VAEs, and rigorously delineate their boundaries. A dual-perspective classification scheme—structured by both architecture and task—is introduced and extended to cross-domain applications (e.g., medical imaging, remote sensing, video). We conduct benchmarking with joint efficiency–performance evaluation and open-source a resource repository featuring 20+ models and standardized evaluation metrics. Key contributions include: (1) the first structured taxonomy for diffusion-based low-level vision; (2) a cross-task transferability analysis framework; and (3) identification of seven critical future research directions—collectively advancing both theoretical foundations and practical deployment of diffusion models in low-level vision.

Comprehensive survey of diffusion models in low-level vision.Evaluation and future directions for diffusion model applications.Exploration of diffusion model frameworks and their correlations.

Latest Papers

What's happening recently
View more

Existing methods struggle to accurately recover the initial noise latent variable from images generated by DDIM, achieving reasonable reconstruction quality but insufficient latent prediction accuracy. This work proposes a hybrid inversion approach that first employs gradient descent for direct inversion and subsequently refines the estimate through fixed-point iteration to more precisely recover the initial latent variable. The study introduces, innovatively, a “self-interpolation test” as a novel evaluation metric to comprehensively assess latent prediction fidelity. Experimental results demonstrate that the proposed method significantly improves both latent prediction accuracy and image reconstruction quality across three benchmark datasets, consistently outperforming existing approaches in self-interpolation test performance.

DDIM inversiondiffusion modelsimage generation inversion

This work addresses the limitations of latent diffusion models, which suffer from detail loss and misalignment between representation and generation objectives due to fixed visual encoders. To overcome these issues, the authors propose an end-to-end pixel-space diffusion Transformer framework that operates directly in the pixel domain without relying on VAE compression. The approach integrates a continuous generation mechanism within a unified multimodal Transformer architecture, sharing a common token space for both images and text. By carefully optimizing noise scheduling, loss weighting, and model scaling strategies, the method significantly enhances fine-grained detail fidelity in high-resolution image synthesis. This paradigm offers a promising direction toward building integrated multimodal vision foundation models capable of both generative and perceptual tasks.

diffusion transformersend-to-end optimizationhigh-fidelity generation

This work addresses the structural ambiguity and training complexity in single-image 3D geometry reconstruction caused by reliance on latent-space compression or hybrid architectures. We propose an extremely minimalist pixel-space diffusion Transformer that trains an end-to-end diffusion model directly on raw 3D point patches, eliminating implicit encoding, complex loss functions, and the need for a point-patch tokenizer. Leveraging pretrained DINOv2 image features as conditioning guidance for geometry generation and built upon a standard ViT backbone, our method achieves the first purely pixel-space diffusion-based geometric reconstruction. It significantly enhances geometric sharpness and robustness—particularly in transparent and highly ambiguous regions—and outperforms existing implicit diffusion approaches.

3D reconstructiondiffusion modelslatent space

This work addresses the high inference latency and computational cost of diffusion models arising from their iterative denoising process. The authors propose a timestep-aware dynamic inference acceleration method that learns dedicated masks for each denoising step to dynamically skip redundant network blocks and reuse intermediate features, thereby reducing computation. To mitigate the high memory overhead of global backpropagation, mask optimization is performed independently per timestep. Stability is further enhanced through timestep-aware loss scaling and a knowledge-guided mask refinement strategy. The approach achieves significant inference speedups across diverse architectures—including DDPM, LDM, DiT, and PixArt—while preserving generation quality.

computational efficiencyDiffusion Probabilistic Modelsimage generation

This work investigates how to construct a diffusion-friendly latent space to enhance generation quality, moving beyond the sole optimization of reconstruction fidelity. The authors systematically evaluate diverse visual tokenizer architectures, regularization strategies, and latent configurations across multiple diffusion backbones. They introduce a novel metric, Velocity Irreducible Variance (VIV), to quantify velocity ambiguity in the latent space arising from trajectory intersections. Experimental results demonstrate that VIV serves as a robust predictor of generation quality, consistently outperforming other latent-space attributes across various settings. The study further uncovers several key characteristics of latent representations that exhibit strong generalization capabilities, offering actionable insights for designing better latent spaces tailored to diffusion models.

diffusion modelsgeneration qualitylatent diffusability

Hot Scholars

HL

Haotong Lin

Zhejiang university
Computer Vision and Graphics
SP

Sida Peng

Zhejiang University
Computer VisionComputer Graphics
GX

Gangwei Xu

Huazhong University of Science and Technology
Computer VisionDeep LearningStereo Matching