🤖 AI Summary
This study addresses the limited reliability and constrained exploration inherent in existing automated red teaming approaches for evaluating the safety of text-to-image models. We propose an attack framework based on iterative strategy evolution that evolves universal attack strategies rather than rewriting individual prompts, thereby significantly enhancing exploration efficiency and cross-model transferability. Furthermore, we define category-specific rigorous success criteria and employ human annotations to calibrate vision-language model (VLM) judges, ensuring evaluation reliability. Experimental results demonstrate that our method achieves a human-verified attack success rate of up to 13% on mainstream systems such as DALL-E 3, substantially outperforming conventional baselines.
📝 Abstract
On modern production text-to-image systems, successful policy violations are rare, and previously effective human-written seeds are often patched out. Current automated red-teamers are poorly matched to this regime in two ways: unreliable success measurement and poor exploration. First, we find that judges widely used in prior T2I red-teaming work are unreliable under vague unsafe-content targets: they either miss true violations or reward benign borderline images on hardened APIs. We therefore define strict category-specific success criteria and calibrate strong VLM judges against human labels. Second, we show that broadly used prompt-modification pipelines do not solve the exploration problem: on harder guardrail settings they remain tied to seed prompts, fail to transfer, or cannot bootstrap positive examples. We introduce RISE, which evolves reusable strategies used to generate prompts rather than rewriting them one by one. The best discovered strategies are then reused to generate attacks across new scenarios. On DALL-E 3, Nano Banana 2 (Google) and GPT-Image-2, RISE reaches up to 13% human-verified ASR; under the same calibrated evaluation, prior methods with reported ASR as high as roughly 30% fall to near zero.