Red-TTT: Test-Time Training for Automated Jailbreaking Large Language Models

📅 2026-10-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limited adaptability and low success rates in large language model (LLM) red teaming caused by frozen weights. To overcome this, we propose Red-TTT, a method that incorporates test-time training to dynamically update attacker parameters via policy gradient during the attack process. Rather than relying solely on in-context signals, Red-TTT consolidates exploration signals directly into the model weights. Furthermore, the training objective is tailored for red teaming by evaluating performance against the best sampled response, requiring only sampling-level access for seamless integration into existing pipelines. Experimental results demonstrate that, under a 120-sample budget, Red-TTT increases the average attack success rate from 55.9% to 72.4%, comprehensively outperforming baselines and successfully eliciting multiple challenging harmful behaviors.
📝 Abstract
Large language models remain vulnerable to jailbreaks, and automated red teaming is the standard way to find jailbreaks in large language models at scale. Current methods either draw more samples at test time through search, rewriting, and tree expansion, or train a stronger attacker offline with reinforcement learning. Both share a limitation: once an attack on a specific target behavior begins, the attacker's weights are frozen. Any signal it gathers about the behavior stays in its context window and is discarded afterward. The attacker never adapts its proposal distribution mid-attack, so success depends almost entirely on the sampling budget, and under a budget affordable at scale, many behaviors remain unbroken. We propose Red-TTT, which updates the attacker's parameters during the attack on each behavior. At each round, the attacker samples a group of candidates, scores them against the victim's replies, and takes a policy-gradient step before drawing the next group, so what it discovers about the current victim is consolidated into weights rather than accumulated as context. We also adapt the training objective to red teaming, where success is judged by the single best sample rather than the average. Red-TTT requires only sampling access to the victim and integrates into existing attack pipelines with no other changes. Against the Best-of-N baseline, Red-TTT raises attack success rate from 55.9\% to 72.4\% on average at a budget of 120 samples, improving over the baseline in every configuration and cracking many behaviors previous method cannot. The code is available at https://github.com/SaFo-Lab/Red-TTT
Problem

Research questions and friction points this paper is trying to address.

automated red teaming
jailbreak
large language models
test-time training
safety evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Test-Time Training
Automated Jailbreaking
Policy Gradient
Red Teaming
Large Language Models
🔎 Similar Papers