Score
Design and implement decoding or synthesis models that, given a scene representation (image, latent embedding, or geometric proxy) and explicit medium controls (e.g., density, scattering, absorption, flow), add, remove, or transform physically plausible participating-medium effects (fog, smoke, water, etc.) while preserving the original scene content; this includes building medium-conditioned decoders and physically based degradation/synthesis pipelines that enforce accurate transport and appearance constraints.
This work addresses the scarcity of real-world training data in underwater computer vision by proposing a two-stage latent diffusion generation framework that fully disentangles scene content from water medium effects in the latent space for the first time. The method first synthesizes degradation-free latent representations of underwater scenes and then applies a physically accurate underwater optical degradation model. Through fine-tuning of the U-Net architecture and a conditional decoding mechanism, the approach enables independent control over image content and underwater degradations. This allows the generation of a large-scale synthetic underwater dataset that preserves both diversity and photorealism, significantly improving performance on downstream tasks such as image restoration and semantic segmentation.
This work addresses the challenges of low fidelity and semantic distortion in e-commerce product image background replacement. We propose an end-to-end recontextualization framework built upon text-to-image diffusion models. Methodologically, we introduce the first integrated data synthesis pipeline combining image-to-video diffusion, intelligent inpainting/outpainting, and negative-sample augmentation, alongside a product representation disentanglement mechanism to jointly optimize structural consistency and attribute fidelity. Experiments on the ABO and proprietary e-commerce datasets demonstrate substantial improvements: FID decreases by 32%, CLIP-Score increases by 18%, and human evaluations show 41% and 53% gains in realism and product consistency, respectively—outperforming state-of-the-art methods. Our core contributions are: (1) the first controllable generation paradigm tailored for product recontextualization; (2) a disentangled product representation learning mechanism; and (3) a multi-stage synthesis strategy that jointly ensures photorealism and semantic consistency.
To address low visual fidelity (e.g., wear, aging, weathering) and geometric inconsistency of physical materials across multi-view observations, this paper proposes a fine-tuning-free differentiable inverse rendering framework. Methodologically, it integrates pre-trained text-to-image diffusion models (e.g., Stable Diffusion) with multi-view differentiable rendering, leveraging UV-space-consistent noise initialization and projection-constrained attention to enforce cross-view geometric and appearance alignment. Coupled with text-conditioned guidance and joint backpropagation over PBR parameters, the framework directly optimizes physically consistent 2D texture maps—including albedo, normal, and roughness. This work is the first to seamlessly embed diffusion priors into the inverse rendering pipeline while preserving material physical differentiability and enabling artist-friendly interactive editing. As a result, it significantly reduces the cost of producing high-fidelity, physically based rendering (PBR) materials.
This paper addresses the problem of generating dynamic visual illusion images. We propose a zero-shot diffusion sampling framework that enables perceptually controllable, text-guided synthesis via image component decomposition—specifically into frequency-domain, luminance/chrominance, and motion blur subspaces. Our method integrates multi-scale frequency decomposition, component-conditioned sampling, and composite noise estimation. Key contributions include: (1) the first noise-estimation factorization fusion mechanism, enabling joint modeling of heterogeneous component-wise conditions; (2) unified generation of appearance variations governed by spatial distance, illumination, and motion blur dependencies; and (3) component-level inverse editing and controllable re-synthesis of real-world images. Crucially, our approach achieves high fidelity and photorealism without compromising sharpness or perceptual plausibility, thereby extending the paradigm of controllable spatial-perception illusion generation.
Conditional image synthesis with diffusion models suffers from a lack of systematic understanding due to architectural complexity, task heterogeneity, and diverse conditional mechanisms. Method: This paper introduces the first unified taxonomy for conditional diffusion modeling, categorizing approaches by *where* conditioning is injected—either into the denoising network architecture or the sampling process—and formalizes three paradigmatic stages: training, reuse, and specialization. It further classifies six mainstream sampling-time conditioning strategies. Contributions: Based on a structured analysis of over 100 works, the paper establishes a comprehensive knowledge framework and open-sources an authoritative resource repository (GitHub Awesome-Conditional-Diffusion-Models). It identifies persistent bottlenecks—including limited generalization, inefficient inference, and coarse-grained control—and proposes principled directions toward scalable, modular, and fine-grained conditional modeling.
This work addresses the challenge of achieving precise and controllable image generation with diffusion models in the absence of large-scale annotated data. The authors propose a self-conditioning mechanism leveraging pretrained self-supervised representations, which identifies semantic directions in the representation space to guide the diffusion process without requiring labeled conditions. This approach not only enhances unconditional generation quality but also constructs a smooth and disentangled controllable generation space. Experimental results demonstrate that the proposed method achieves superior performance in image generation and editing tasks, excelling in controllability, smoothness, and disentanglement compared to existing alternatives.
This study addresses the degradation of high-fidelity details in image-to-3D generation caused by global encoding compression. To overcome this limitation, we propose BTC3D, a training-free inference framework that first reveals the additivity of diffusion model features. By designing hybrid tile embeddings coupled with a dynamic conditioning scheduling mechanism, our method precisely extracts and fuses local high-frequency signals, enhancing detail preservation without requiring retraining. Experimental results demonstrate that BTC3D significantly improves texture quality and visual fidelity while maintaining global structural consistency. Furthermore, it can be seamlessly integrated into existing generative pipelines.
This work addresses the degradation of image visibility and multi-view consistency caused by haze, which severely compromises novel view synthesis quality. To tackle this challenge, the authors propose a multi-stage optimization framework that sequentially integrates image restoration, dehazing, enhancement via multimodal large language models (MLLMs), and joint optimization of 3D Gaussian Splatting (3DGS) with Markov Chain Monte Carlo (MCMC), followed by averaging across multiple refinement rounds to simultaneously enhance visibility and preserve cross-view scene consistency. By synergistically combining generative priors with geometric optimization, the method achieved first place among 14 teams in Track 2 of the NTIRE 2026 3DRR Challenge, demonstrating superior quantitative performance and visual quality on the official benchmark.
This work addresses the challenge of effectively leveraging self-supervised high-dimensional features—such as those from DINO—for controllable video generation within pretrained video diffusion models, while disentangling appearance from scene attributes like semantics and geometry that should be preserved. To this end, the authors propose a lightweight conditional injection architecture combined with a feature disentanglement training strategy, introducing self-supervised features as a general-purpose control signal into video diffusion models for the first time. The approach enables appearance-editing tasks such as style transfer and relighting, significantly enhances generation controllability at low spatial resolutions, and successfully achieves high-quality video domain transfer and 3D-to-video synthesis.
Existing diffusion-based object insertion methods treat the task solely as 2D image inpainting, lacking explicit control over the 3D pose of inserted objects. This work proposes DIRECT, a novel framework that, for the first time, integrates user-controllable 3D proxies with 2D diffusion generation. By decoupling appearance, geometry, and contextual guidance signals and injecting them into separate pathways, DIRECT effectively mitigates feature entanglement. The approach enables precise 3D pose control and scene-adaptive placement while preserving the visual fidelity of reference objects. Experimental results demonstrate that DIRECT outperforms existing methods in both geometric controllability and visual quality, supporting high-fidelity object insertion under interactive 3D pose adjustments.