🤖 AI Summary
This work proposes a scalable evaluation framework to effectively assess the perceptual realism of synthetic images generated under adverse environmental conditions such as fog, rain, snow, and nighttime. The framework innovatively integrates a visual-language model (VLM) jury for perceptual realism scoring with distributional similarity analysis in embedding space, enabling, for the first time, a unified and efficient evaluation of both generative and rule-based image enhancement methods, with real images serving as the performance upper bound. Experimental results demonstrate that generative AI approaches significantly outperform rule-based methods, with the best-performing model achieving an acceptance rate approximately 3.6 times higher than that of rule-based counterparts; under most conditions, its synthetic images attain or even surpass the realism of real images.
📝 Abstract
Evaluation of AI systems often requires synthetic test cases, particularly for rare or safety-critical conditions that are difficult to observe in operational data. Generative AI offers a promising approach for producing such data through controllable image editing, but its usefulness depends on whether the resulting images are sufficiently realistic to support meaningful evaluation.
We present a scalable framework for assessing the realism of synthetic image-editing methods and apply it to the task of adding environmental conditions-fog, rain, snow, and nighttime-to car-mounted camera images. Using 40 clear-day images, we compare rule-based augmentation libraries with generative AI image-editing models. Realism is evaluated using two complementary automated metrics: a vision-language model (VLM) jury for perceptual realism assessment, and embedding-based distributional analysis to measure similarity to genuine adverse-condition imagery.
Generative AI methods substantially outperform rule-based approaches, with the best generative method achieving approximately 3.6 times the acceptance rate of the best rule-based method. Performance varies across conditions: fog proves easiest to simulate, while nighttime transformations remain challenging. Notably, the VLM jury assigns imperfect acceptance even to real adverse-condition imagery, establishing practical ceilings against which synthetic methods can be judged. By this standard, leading generative methods match or exceed real-image performance for most conditions.
These results suggest that modern generative image-editing models can enable scalable generation of realistic adverse-condition imagery for evaluation pipelines. Our framework therefore provides a practical approach for scalable realism evaluation, though validation against human studies remains an important direction for future work.