Score
Design and implement methods that generate image perturbations by manipulating scene illumination, including spatially varying relighting and intensity modulation across regions. Build and analyze illumination control schemes and image-perturbation techniques that allow per-region strength adjustment while maintaining realistic lighting appearance and balancing perturbation strength against perceptual image quality.
This work proposes an attribute-token-based method for image relighting, formulating illumination editing as a conditional image generation task. By introducing learnable attribute tokens, the approach disentangles and implicitly encodes multidimensional lighting parameters—including intensity, color, ambient light, diffuse components, and 3D light source positions—without requiring inverse rendering supervision. This enables implicit modeling of the complex interactions among lighting, geometry, material properties, and occlusion. Leveraging a diffusion-based generative architecture, large-scale synthetic data pretraining, and fine-tuning on real images, the method achieves state-of-the-art performance on both synthetic and real-world datasets. It demonstrates exceptional realism and controllability in challenging scenarios such as placing light sources inside objects or relighting transparent materials.
Existing single-image relighting methods struggle to achieve explicit, disentangled control over lighting parameters—or rely on multi-view inputs and geometric reconstruction. To address this, we propose the first single-image relighting method that embeds the linear photometric properties of light into a diffusion model fine-tuning paradigm. Our approach enables independent, fine-grained adjustment of target light intensity, correlated color temperature, and ambient illumination. Built upon lightweight fine-tuning of Stable Diffusion, it introduces a physics-inspired RGB-scale chromaticity–intensity perturbation mechanism and employs a paired distillation training strategy jointly driven by real image pairs and large-scale synthetic renderings. On single-image relighting, our method significantly outperforms state-of-the-art approaches: user studies show a 37% increase in preference rate, while maintaining real-time interactivity, high photometric consistency, and superior texture fidelity.
This work addresses the challenge of physically plausible relighting of single real-world photographs, where explicit illumination parameters are absent and precise lighting control is thus infeasible. We propose the first method enabling physically controllable relighting from a single unstructured image. Our approach reconstructs a color-coded 3D mesh via monocular depth estimation and intrinsic image decomposition, then integrates path tracing—ensuring physical accuracy—with feedforward neural rendering—guaranteeing visual fidelity—to support interactive, 3D-space lighting parameter manipulation. Crucially, we introduce a self-supervised learning framework that jointly optimizes geometry, material, and lighting end-to-end, without requiring paired data or ground-truth annotations. Trained exclusively on unlabeled real images, our method produces results that are both physically consistent and photorealistic. This represents the first successful extension of explicit, physics-based lighting control—previously confined to controlled 3D graphics settings—to the challenging single-image relighting task.
Existing illumination editing methods struggle to simultaneously achieve customizable lighting control and content fidelity, particularly in cross-image complex illumination transfer tasks. To address this, we propose a generative disentanglement framework that, for the first time, achieves complete separation of content and illumination features in real-world scenes. We construct a million-scale image–content–illumination triplet dataset and perform end-to-end training on a fine-tuned diffusion model integrated with the IC-Light architecture, conditioned on reference illumination maps. Our approach enables high-fidelity, highly flexible illumination transfer, significantly improving cross-domain illumination harmony and editing naturalness. Quantitative and qualitative evaluations demonstrate that our method surpasses state-of-the-art approaches across multiple metrics—including LPIPS, SSIM, and user studies—while enabling unprecedented control over lighting attributes without compromising structural or textural integrity. This work establishes a new paradigm for illumination editing and harmonization, advancing both theoretical understanding and practical applicability in photorealistic image synthesis.
Existing 3D neural radiance field (NeRF) stylization methods lack fine-grained controllability in color fidelity, stylistic scale, spatial selectivity, and depth awareness. To address these limitations, we propose the first four-dimensionally aware controllable NeRF stylization framework, enabling simultaneous color preservation, adjustable stylistic pattern scaling, mask-guided local stylization, and depth-aware regularization. Our approach introduces a multi-objective differentiable loss function and an end-to-end optimization strategy to support seamless multi-style fusion and user-customized generation. Extensive evaluations on multiple real-world scene datasets demonstrate high-fidelity rendering, real-time interactivity, and professional-grade artistic editing capabilities. The framework significantly advances controllability and practicality in 3D NeRF stylization, establishing new benchmarks for editable, geometry-aware neural rendering.
Existing single-image relighting methods suffer from limited lighting control, error accumulation in cascaded pipelines, or the need for per-image optimization. This work proposes a feed-forward relighting framework that, for the first time, integrates intrinsic image decomposition with neural rendering. By sharing intrinsic cues—such as albedo, diffuse shading, and non-diffuse residuals—it bridges physics-based rendering and learned synthesis, enabling arbitrary physically based rendering (PBR)-compatible lighting edits. The approach leverages a Transformer-based neural renderer, path-tracing-guided coarse 3D reconstruction, and pixel-wise affine modulation to achieve high-quality relighting at unprecedented speed (<0.1 seconds per image), while preserving physical plausibility and fine details. The method attains state-of-the-art performance across standard benchmarks.
This work proposes a novel perturbation-based explainable artificial intelligence (XAI) method that integrates generative image inpainting into the LIME framework to address the limitations of conventional perturbation techniques, which often produce artifacts and distributional shifts that degrade explanation quality. By leveraging generative inpainting, the approach synthesizes perturbed samples that adhere closely to the original data distribution and exhibit high visual fidelity. This effectively eliminates visible artificial traces while preserving local faithfulness, thereby substantially enhancing both the accuracy and credibility of model explanations. The method establishes a more reliable foundation for visual XAI by ensuring that generated perturbations are perceptually realistic and semantically consistent with the input data.
Existing diffusion models struggle to simultaneously achieve photorealism, spatial lighting control accuracy, and multi-view consistency in indoor scene relighting, often compromising generative priors when specifying 3D light source positions. To address this, we propose Lume-Palette, a framework that decouples relighting into two stages: illumination distillation and illumination projection. The first stage extracts a “lighting palette” from a pre-trained diffusion model, preserving realistic material–illumination interactions, while the second stage explicitly maps target spatial illumination onto a coarse 3D geometry. By disentangling semantic lighting priors from spatial control and incorporating an asymmetric multi-view conditioning strategy, our method achieves high-fidelity, spatially accurate, and multi-view consistent relighting results on both synthetic and real-world indoor scenes.
本文通过引入材料解耦的光照表示法Lumi Map,解决了手绘草图控制图像重新打光时精度和可控性不足的问题。
This work addresses the limitations of virtual production, where LED wall backgrounds are tightly coupled with lighting, restricting relighting flexibility in post-production, and where traditional inverse rendering based on environment maps lacks accuracy under near-field, high-resolution illumination. To overcome these challenges, the authors propose a relightable Gaussian splatting framework tailored for virtual production that dispenses with environment maps entirely. Instead, background textures are sampled directly in image space using Gaussian primitives modulated by UV coordinates, intensity, and resolution, implicitly modeling reflections and refractions. Leveraging known background information, relighting is guided and reduced to an image editing task. By decomposing appearance and lighting from multi-illumination real-world data, the method achieves high-quality reconstruction and controllable relighting at ~35 FPS with under two hours of training and less than 5 GB of GPU memory, while supporting outputs such as depth, light intensity, color, and unlit renderings.