🤖 AI Summary
This study addresses the difficulty vision-language models (VLMs) face in recognizing text and contours concealed within diffusion-generated images. To overcome the parameter- and viewpoint-dependence of conventional transformations, this work proposes directly recovering the control fields employed during image generation. Specifically, a lightweight U-Net is utilized to predict grayscale control fields, which are then integrated with Qwen2.5-VL to enable hidden content recognition. Furthermore, the FreqBlind benchmark is constructed for systematic evaluation. Experimental results demonstrate that the proposed approach achieves 60.2% accuracy in contour recognition, surpassing the baseline by 26.9 percentage points. Additionally, inference incurs only 7.4 ms of extra latency while maintaining robustness against compression noise.
📝 Abstract
Spatially conditioned diffusion models can embed words and contours in natural-looking images, but vision-language models (VLMs) may fail to recognize the hidden content. Transformation-based recovery depends on parameter and view selection. To evaluate hidden-content recovery and recognition, we construct FreqBlind, a 6,000-image benchmark spanning contours, real words and non-words across three conditioning strengths. The evaluated transformation-based methods show limited recognition of contour patterns and weakly conditioned hidden content. To address this limitation, we propose ControlTrace to recover the grayscale control field used during generation. An 8.4M-parameter U-Net predicts this field from the carrier image, and a VLM then identifies its content. With Qwen2.5-VL-7B-Instruct, ControlTrace achieves 60.2% open-ended contour recognition accuracy across the three conditioning strengths, exceeding the best of the three evaluated prior methods by 26.9 percentage points. On an A100 GPU, the complete pipeline adds only 7.4 ms (5.3%) to direct VLM inference. Recovered fields have lower pixel errors and higher structural similarity than the evaluated transformation views. Across four evaluated VLMs, ControlTrace retains its overall contour recognition advantage. Recognition remains stable under the tested JPEG compression, Gaussian noise and downsampling. These results support control-field recovery for hidden-content recognition in the evaluated setting.