Score
Design and build algorithms and pipelines that estimate scene illumination (including illuminant spectra), perform photometric normalization (white balance, exposure matching, shadow and reflection compensation) and harmonize lighting across images, frames, and composites. Implement global-prior and global-to-local, coarse-to-fine optimization and spectrum-aware models for standard and multispectral data to align exposures and illuminants while preserving structural consistency and avoiding excessive noise amplification.
Existing spectral response modeling for digital cameras is typically confined to isolated components, lacking an end-to-end, physically consistent description from illumination input to pixel intensity output—thus limiting color fidelity and spectral accuracy. This paper introduces the first full-chain, physics-driven end-to-end spectral–color joint modeling framework. It unifies the coupled effects of optical system transmission, sensor quantum efficiency, color filter array (CFA) spectral transmittance, and nonlinear pixel response. The model integrates empirically measured RGB camera spectral responses with data-driven nonlinear mapping correction. Evaluated under multiple illuminants, it achieves superior color reproduction (mean ΔE < 1.2) and significantly improved spectral reconstruction fidelity (37% reduction in RMSE). Validated across machine vision, remote sensing, and computational spectral imaging applications, this work bridges a critical theoretical and practical gap in end-to-end camera spectral response modeling.
This work addresses color correction inaccuracies in digital cameras under complex spectral and high-chromaticity LED illumination, which arise from the nonlinear relationship between sensor responses and the CIE XYZ color space. To mitigate this, the authors propose an illumination-adaptive 3D lookup table framework, termed C²LUT, that integrates chromaticity-aware illumination representation with nonlinear color transformation. The method employs Tucker tensor decomposition to compress the lookup table, achieving a favorable trade-off between colorimetric accuracy and hardware deployment efficiency. Evaluated on a large-scale dataset comprising 1,473 spectral illuminants, C²LUT demonstrates consistent performance gains across multiple cameras, diverse lighting conditions, and real-world imagery—reducing CIE ΔE₀₀ error by up to 20% and angular error by as much as 18%, while adhering to the computational constraints of modern image signal processors (ISPs).
Existing white balance (WB) correction methods suffer from low accuracy in multi-illuminant scenes, and conventional linear fusion approaches fail to ensure cross-region color consistency. Method: This paper proposes the first Transformer-based image fusion framework tailored for multi-illuminant WB correction. It performs end-to-end nonlinear fusion of multiple preset-WB images of the same scene directly in the sRGB domain—replacing traditional linear weighting—and introduces a lightweight spatial-illumination-aware Transformer module to jointly model local chromatic bias and global illumination interactions. Contribution/Results: To enable training and evaluation, we introduce ML-WB, the first large-scale multi-illuminant WB dataset comprising over 16,000 images, covering five preset WB settings under complex mixed-illumination conditions. Experiments on ML-WB demonstrate that our method achieves up to 100% improvement over state-of-the-art methods, significantly enhancing both color consistency and cross-scene generalization capability.
To address the limited color correction accuracy of mobile device cameras constrained by single-modality RGB input, this paper proposes an end-to-end jointly optimized framework that fuses high-resolution RGB with low-resolution multispectral sensor data. Unlike conventional approaches relying on hand-crafted priors or feature concatenation, our method preserves the full multispectral information flow throughout the pipeline and unifies the modeling of sensor response, spectral reconstruction, and color mapping—enabling seamless integration with state-of-the-art image architectures. Trained on a custom multispectral rendering dataset, the model achieves computational efficiency while significantly improving cross-device robustness. Experiments demonstrate a reduction of up to 50% in mean color error (ΔE₀₀) compared to the best RGB-only baseline, outperforming methods using multispectral priors alone, and exhibiting superior stability across diverse hardware spectral responses.
Existing shadow removal methods suffer from limited performance in multi-light-source environments due to misalignment between physical priors and data-driven features. This work proposes a dual-level prior alignment framework that first employs Physically Aligned Normalization (PAN) to achieve closed-form illumination correction, followed by a Geometry-Semantic Rectification Attention (GSRA) mechanism that fuses depth geometry with DINO-v2 semantic features to enhance cross-modal consistency. For the first time, this approach enables synergistic alignment between physical priors and semantic-geometric representations, significantly improving robustness and generalization across scenarios ranging from single to complex multi-light settings. The method achieves superior performance on multiple real-world shadow datasets while maintaining lower computational complexity compared to existing approaches.
This work addresses the challenge that existing methods struggle to effectively exploit spectral information in multispectral images under diverse illumination and cross-sensor conditions, leading to limited performance in illuminant spectral estimation. To overcome this, we propose a novel deep learning framework that integrates illumination priors through a spatial-spectral feature extraction module coupled with a spectral attention mechanism to enhance responses in critical spectral channels. Additionally, we introduce a training-free spectral-domain transformation strategy that enables efficient transfer of illuminant spectra from high-dimensional multispectral sensors to low-dimensional camera sensors. Experiments on a newly constructed real-world multispectral dataset demonstrate that our method significantly outperforms state-of-the-art approaches, achieving highly accurate and robust illuminant spectral estimation.
该论文提出了一种基于几何驱动的阴影调和方法,通过估计每个像素的增益场并进行RGB通道均匀乘法处理,解决了面部合成中光照不一致的问题。
This study addresses the performance bottleneck in illumination estimation models caused by the scarcity of real-world annotated data. To overcome this limitation, we propose a reusable physics-based synthetic data pipeline that leverages physically based rendering to generate pixel-level light source annotations and dense chromaticity maps. Using this pipeline, we construct a large-scale synthetic dataset for pre-training both single- and multi-illuminant estimation models. This work effectively circumvents the challenge of acquiring high-quality annotations and substantially enhances generalization capabilities in few-shot scenarios. Experimental results demonstrate that pre-training with the proposed synthetic data reduces estimation errors by 28% for single-illuminant and 57% for multi-illuminant tasks, respectively, validating the superiority of this synthetic data-driven paradigm.
Current color constancy methods suffer from limited generalization across cameras due to overfitting to the specific color response characteristics of training cameras. This work proposes VLM-CC, a novel framework that formulates color constancy as an iterative optimization process driven by a vision-language model (VLM). Specifically, a LoRA-finetuned VLM provides semantic-aware perceptual assessments of chromatic bias in images, which, combined with pseudo-sRGB conversion and residual illumination direction mapping, progressively refines white balance estimates. Departing from conventional RGB regression paradigms, the method leverages semantic feedback to guide optimization, achieving state-of-the-art cross-camera robustness and significantly outperforming existing approaches on multiple benchmark datasets.
Existing methods for video portrait relighting often suffer from temporal flickering and incoherence due to the lack of paired real-world data. To address this, this work proposes a deflickering illumination model that generates temporally stable paired videos for training. By integrating a hybrid strategy combining real and synthetic videos, we develop a video diffusion model enhanced with an asymmetric Alpha mask conditioning mechanism to improve boundary sharpness and temporal consistency. The proposed approach maintains strong relighting capabilities while significantly outperforming current image- and video-level methods, achieving notable improvements in temporal stability, visual naturalness, boundary quality, and physically plausible lighting effects.