FinRT: Distilling Adaptive Red-Teaming Strategies into Reusable Adversarial Generators in Consumer Finance

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the safety failures in consumer finance large language models triggered by covert queries, as well as the challenge faced by existing red teaming methods in balancing attack effectiveness with computational cost. To this end, we propose the first structured adversarial generation framework. By integrating adaptive red teaming strategies, adversarial prompt generation, and model distillation, our approach transforms one-time searches into reusable generators, achieving joint optimization of coverage, severity, and diversity. Experimental results demonstrate that the proposed method increases the attack success rate to 32.9% and improves maximum adversarial severity by 33%. Furthermore, it preserves semantic diversity while exhibiting strong cross-model transferability.
📝 Abstract
In regulated industries like consumer finance, seemingly harmless user queries can exploit large language model vulnerabilities, triggering safety failures and pushing responses dangerously close to policy limits. Existing automated red-teaming methods trade off attack effectiveness against generation cost, while treating coverage, severity, and diversity as incidental rather than joint objectives. We introduce FinRT, a structured framework that builds reusable adversarial prompt generators from adaptive red-teaming strategies. Across the six victim models in consumer finance, FinRT substantially outperforms adaptive search baselines while amortizing target-facing attack generation into a reusable generator. FinRT nearly doubles the attack success rate over the adaptive baseline Rainbow Teaming (32.9% vs. 17.2%), increases maximum adversarial severity by 33%, and preserves comparable intra-policy-domain semantic diversity to iterative search methods. Our method achieves high cross-model transferability while exhibiting distinct victim-family specialization patterns.
Problem

Research questions and friction points this paper is trying to address.

red-teaming
large language models
consumer finance
adversarial prompts
safety failures
Innovation

Methods, ideas, or system contributions that make the work stand out.

Red-teaming
Adversarial prompt generation
Knowledge distillation
LLM safety
Consumer finance