Evaluating Language Model Safety Across Long Adversarial Conversations

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of existing single-turn safety evaluations, which fail to capture the progressive degradation of model safety under sustained adversarial interactions in long conversations. To investigate this, we employ large language models to simulate persistent malicious users and conduct multi-turn adversarial testing on open-source instruction-tuned models, utilizing automated classifiers for turn-level safety annotation. This work provides the first empirical demonstration of risk accumulation in extended dialogues: while initial safety rates range from 85% to 100%, they precipitously decline to 15%–44% by the 101st turn. By offering a proof-of-concept that strong single-turn safety does not guarantee long-term robustness, this research transcends conventional short-interaction evaluation paradigms and underscores the necessity of long-horizon safety assessments.
📝 Abstract
Conversational safety evaluations often test language models with a single harmful prompt, even though real-world systems interact with users through long, adaptive conversations. This study examines whether models continue to respond safely when an adversarial user persists across multiple turns. We evaluate three open-weight, instruction-tuned models on two harmful prompts across different conversation lengths and random seeds. In each setting, a second language model acts as a persistent adversarial user, while a safety classifier labels every response as safe or unsafe. Across all model-prompt combinations, first-turn safe-response rates ranged from 85% to 100%. By depth 11, they dropped to 38-61%, and by depth 101, to 15-44%. This decline appeared across models and continued well beyond the short interactions typically used in multi-turn safety evaluations. These results provide proof-of-concept evidence that strong single-turn safety does not necessarily persist during sustained adversarial interaction. They highlight the need for long-horizon evaluations and conversation-level safeguards that account for risk accumulating across turns.
Problem

Research questions and friction points this paper is trying to address.

language model safety
adversarial conversations
multi-turn evaluation
long-horizon interactions
Innovation

Methods, ideas, or system contributions that make the work stand out.

Adversarial Conversations
Language Model Safety
Multi-turn Evaluation
Safety Classifier
Risk Accumulation