🤖 AI Summary
This study addresses the performance bottleneck in illumination estimation models caused by the scarcity of real-world annotated data. To overcome this limitation, we propose a reusable physics-based synthetic data pipeline that leverages physically based rendering to generate pixel-level light source annotations and dense chromaticity maps. Using this pipeline, we construct a large-scale synthetic dataset for pre-training both single- and multi-illuminant estimation models. This work effectively circumvents the challenge of acquiring high-quality annotations and substantially enhances generalization capabilities in few-shot scenarios. Experimental results demonstrate that pre-training with the proposed synthetic data reduces estimation errors by 28% for single-illuminant and 57% for multi-illuminant tasks, respectively, validating the superiority of this synthetic data-driven paradigm.
📝 Abstract
Illuminant estimation is a fundamental problem in computational photography, as it enables the correction of color shifts induced by varying lighting conditions. While learning-based methods have demonstrated strong performance, their progress is hindered by the limited availability of large-scale datasets with accurate illuminant ground-truth. In this work, we propose a general and reusable pipeline to derive dense illuminant chromaticity maps from physically based 3D-rendered scenes. By repurposing an existing 3D scene collection, our approach enables the systematic generation of pixel-wise illuminant annotations under controlled lighting conditions, effectively lowering the barrier to data acquisition for learning-based illuminant estimation. Using this pipeline, we generate a large-scale synthetic set of 74,321 images, which we employ for pre-training both single- and multi-illuminant estimation models. Extensive experiments with state-of-the-art architectures show that synthetic pre-training consistently improves performance, with gains of up to 28% for single-illuminant estimation and up to 57% for multi-illuminant estimation, particularly in data-scarce regimes. These findings demonstrate that synthetic data generation pipelines offer an effective and scalable solution for the pre-training of illuminant estimation methods.