CART: Closed-Loop Adaptive Red Teaming for Large Language Models

📅 2026-09-23
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
CART框架通过自适应循环测试方法,不断根据测试结果调整策略,以发现大型语言模型中的更多弱点和风险,优于固定提示重放。
📝 Abstract
Automated red teaming often replays a fixed set of prompts, which measures known risks but cannot learn from failures found during testing. We present CART (Closed-Loop Adaptive Red Teaming), a framework that uses each result to guide what it tests next. CART begins with broad risk coverage, follows weaknesses that emerge, keeps new probes diverse, and records the evidence and source of every finding. It separates the Challenger that creates tests, the Target being tested, which may be a text-only model or a bounded tool-using agent, and the Judge that evaluates the results, allowing these roles to be studied independently. Across three evaluation families (Frontier, JAH, and Agentic), CART discovers more failures and higher average risk than static seed replay for every Target with an available baseline. The gains extend to tool-mediated agent tests, suggesting that contextual adaptation can reveal weaknesses that direct prompt replay does not exercise. These results describe what the test policies discover, not how often failures occur in real deployments. We also find that Challenger-Judge choices affect the evidence uncovered, highlighting the need for role separation and independent review. Overall, CART turns red teaming from a one-time checklist into a continuous, adaptive, and auditable search for model and agent weaknesses.
Problem

Research questions and friction points this paper is trying to address.

Automated Red Teaming
Adaptive Testing
Language Models
Risk Assessment
Innovation

Methods, ideas, or system contributions that make the work stand out.

Closed-Loop Adaptive Red Teaming
contextual adaptation
role separation
🔎 Similar Papers
No similar papers found.