RISE: Red-teaming via Iterative Strategy Evolution for Modern Text-to-Image Models

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limited reliability and constrained exploration inherent in existing automated red teaming approaches for evaluating the safety of text-to-image models. We propose an attack framework based on iterative strategy evolution that evolves universal attack strategies rather than rewriting individual prompts, thereby significantly enhancing exploration efficiency and cross-model transferability. Furthermore, we define category-specific rigorous success criteria and employ human annotations to calibrate vision-language model (VLM) judges, ensuring evaluation reliability. Experimental results demonstrate that our method achieves a human-verified attack success rate of up to 13% on mainstream systems such as DALL-E 3, substantially outperforming conventional baselines.
📝 Abstract
On modern production text-to-image systems, successful policy violations are rare, and previously effective human-written seeds are often patched out. Current automated red-teamers are poorly matched to this regime in two ways: unreliable success measurement and poor exploration. First, we find that judges widely used in prior T2I red-teaming work are unreliable under vague unsafe-content targets: they either miss true violations or reward benign borderline images on hardened APIs. We therefore define strict category-specific success criteria and calibrate strong VLM judges against human labels. Second, we show that broadly used prompt-modification pipelines do not solve the exploration problem: on harder guardrail settings they remain tied to seed prompts, fail to transfer, or cannot bootstrap positive examples. We introduce RISE, which evolves reusable strategies used to generate prompts rather than rewriting them one by one. The best discovered strategies are then reused to generate attacks across new scenarios. On DALL-E 3, Nano Banana 2 (Google) and GPT-Image-2, RISE reaches up to 13% human-verified ASR; under the same calibrated evaluation, prior methods with reported ASR as high as roughly 30% fall to near zero.
Problem

Research questions and friction points this paper is trying to address.

red-teaming
text-to-image models
automated evaluation
prompt exploration
safety guardrails
Innovation

Methods, ideas, or system contributions that make the work stand out.

Red-teaming
Text-to-Image
Iterative Strategy Evolution
VLM Judges
Prompt Generation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Dmitrii Kharlapenko
Dmitrii Kharlapenko
MSc student, ETH
Mechanistic InterpretabilityAI Alignment
S
Sergei Bratchikov
K
Konstantin Korolev
A
Aleksandr Nikolich