Reflections and Fragments: Securing LLMs Against Sequential Mosaic Attacks

📅 2026-10-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses mosaic attacks against large language models, wherein individually benign single-turn queries become harmful when combined across multiple turns. We present the first formal defense theory for such attacks, proving fundamental limitations imposed by finite context windows and establishing lower bounds on state and query complexity. Building upon these theoretical insights, we propose a defense framework based on an online state machine termed "Sentinel," which theoretically guarantees zero-failure defense while preserving model utility. Furthermore, we construct a constrained multi-turn self-play mechanism integrated with LoRA fine-tuning to optimize the attack-defense equilibrium. Experimental results demonstrate that our approach significantly enhances both attack-induction capability and defensive robustness, exhibiting strong generalization performance on previously unseen adversarial targets.
📝 Abstract
Self-play red-teaming improves language-model safety by pitting attacker and defender roles against each other in a zero-sum game. However, real adversaries increasingly use mosaic attacks: multi-turn sequences whose individual fragments are innocuous in isolation yet assemble into a harmful payload. We develop a theory of mosaic defense that characterizes what is required to prevent such attacks without sacrificing helpfulness. We first show that no fixed bounded window of recent prompts is sufficient in general: safety-relevant information may occur arbitrarily far back in the interaction. We formalize a watchman, an online state mechanism that carries this information forward, and show that under explicit assumptions it enables zero-failure defense with positive benign helpfulness. Under stronger conditions, it is also optimal among zero-failure defenders. An exact watchman may nevertheless require exponentially many states, while exact maliciousness detection can require exponentially many queries in an unstructured black-box model. These state and query lower bounds do not by themselves imply hard learning: the construction underlying the state lower bound is efficiently learnable from labeled examples, whereas certifying worst-case safety can require substantially more information under restricted access. We also show that self-play equilibrium alone does not certify usefulness, motivating a constrained formulation that maximizes worst-case benign helpfulness among zero-failure defenders. Empirically, training role-specific attacker and defender LoRA adapters over frozen LLMs via multi-turn self-play strengthens both roles: attackers become more effective at eliciting harmful responses, while defenders become more robust to attack, with improvements also observed on unseen attack objectives.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
Mosaic Attacks
Multi-turn Safety
Red-teaming
Sequential Attacks
Innovation

Methods, ideas, or system contributions that make the work stand out.

Mosaic Attacks
Watchman Mechanism
Self-play Red-teaming
LoRA Adapters
Zero-failure Defense
🔎 Similar Papers
No similar papers found.