🤖 AI Summary
This study addresses the physical inconsistency arising from the missing radiometric scale when inferring polarization from RGB images. We propose PolarScale, a benchmark that reformulates this unidentifiable radiometric scale into a dataset-conditioned semantic estimation task for the first time. Leveraging trichromatic full-Stokes measurement data, we conduct systematic evaluations using both restoration-based and generation-based backbone networks alongside multidimensional physical metrics. The optimal model achieves a scale error as low as 3.6% and a physical violation rate below 0.25%, substantially enhancing performance in diffuse reflection separation, material segmentation, and glare classification.
📝 Abstract
Polarization imaging provides physical cues beyond intensity imaging but typically requires specialized hardware. Recent methods infer polarization from RGB-like inputs, yet predict only normalized Stokes components or relative descriptors, from which the radiometric scale needed for full Stokes reconstruction has been divided out. We introduce PolarScale, a benchmark that makes this scale an explicit prediction and evaluation target. Built on existing trichromatic full-Stokes measurements, PolarScale takes the per-scene normalized total-intensity image $s_0$ (a scene-referred linear image, not a consumer sRGB photograph) and asks models to predict normalized Stokes components, AoLP/DoLP/DoCP, and a per-scene scale. Because the scale is divided out of the input, it is not physically identifiable; PolarScale therefore evaluates dataset-conditioned semantic scale estimation against a constant-scale control, together with angular, self-consistency, and physical-bound metrics. Across seven restoration-based and generative backbones and three prediction strategies, the strongest restoration models estimate the scale with 3.6-4.3% mean relative error versus 5.7% for the constant control and violate physical bounds on fewer than 0.25% of pixels, whereas two generative baselines collapse to a near-zero scale; explicit descriptor supervision improves descriptor accuracy (23.66 vs. 18.88 dB PSNR for MAE). Predicted full-Stokes representations improve diffuse/specular separation, material segmentation, and glare classification, although in diffuse/specular separation the learned scale performs only on par with the constant control.