🤖 AI Summary
This paper addresses the problem of generating dynamic visual illusion images. We propose a zero-shot diffusion sampling framework that enables perceptually controllable, text-guided synthesis via image component decomposition—specifically into frequency-domain, luminance/chrominance, and motion blur subspaces. Our method integrates multi-scale frequency decomposition, component-conditioned sampling, and composite noise estimation. Key contributions include: (1) the first noise-estimation factorization fusion mechanism, enabling joint modeling of heterogeneous component-wise conditions; (2) unified generation of appearance variations governed by spatial distance, illumination, and motion blur dependencies; and (3) component-level inverse editing and controllable re-synthesis of real-world images. Crucially, our approach achieves high fidelity and photorealism without compromising sharpness or perceptual plausibility, thereby extending the paradigm of controllable spatial-perception illusion generation.
📝 Abstract
Given a factorization of an image into a sum of linear components, we present a zero-shot method to control each individual component through diffusion model sampling. For example, we can decompose an image into low and high spatial frequencies and condition these components on different text prompts. This produces hybrid images, which change appearance depending on viewing distance. By decomposing an image into three frequency subbands, we can generate hybrid images with three prompts. We also use a decomposition into grayscale and color components to produce images whose appearance changes when they are viewed in grayscale, a phenomena that naturally occurs under dim lighting. And we explore a decomposition by a motion blur kernel, which produces images that change appearance under motion blurring. Our method works by denoising with a composite noise estimate, built from the components of noise estimates conditioned on different prompts. We also show that for certain decompositions, our method recovers prior approaches to compositional generation and spatial control. Finally, we show that we can extend our approach to generate hybrid images from real images. We do this by holding one component fixed and generating the remaining components, effectively solving an inverse problem.