Detecting Unseen Jailbreak Sources: A Multi-Source Conformal Detection Perspective

πŸ“… 2026-10-06
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the challenge of detecting unknown attack sources against large language model jailbreak defenses after deployment. We propose Persist-3, an algorithm that formulates unknown-source detection as a sequential testing problem. Specifically, it leverages multi-source conformal prediction to compute p-values relative to known attack distributions and employs persistent statistics to identify out-of-distribution attack streams. Furthermore, we design a detection guarantee mechanism relying solely on a finite observation window and theoretically prove that short windows suffice to efficiently detect well-separated unknown sources. Experiments conducted on MTK and Prompt Guard 2 demonstrate that the proposed method significantly enhances the system’s detection capability and robustness against unseen jailbreak attacks.
πŸ“ Abstract
Jailbreak defenses for large language models are usually calibrated on a fixed set of known attack methods. However, new attack methods keep appearing after deployment, and existing detectors are observed to degrade under the new attacks (Piet et al., 2025). This paper studies how to detect, from a stream of attack prompts, that the prompts come from an attack source outside the known ones. We formulate it as a sequential test whose null hypothesis is that the stream comes from one of the known sources. Methodology-wise, we compute a conformal p-value against each known source, and reject the null only when the evidence against every known source is large. Using these source-wise p-values, we build Persist-3 and analyze its stationary false-alarm probability and detection power. The technical contribution of Persist-3 lies in a detection guarantee that uses only a finite window of recent observations. Our theoretical analysis reveals that its detection probability approaches one exponentially fast as the window grows, and thus a short window suffices to detect unseen sources that are well separated from all known sources. Empirically, we validate Persist-3 on top of two recent jailbreak defenses, MTK (Zhang et al., 2026) and Prompt Guard 2 (Chennabasappa et al., 2025), and observe a significant defense enhancement.
Problem

Research questions and friction points this paper is trying to address.

Jailbreak detection
Unseen attack sources
Large language models
Conformal detection
Sequential testing
Innovation

Methods, ideas, or system contributions that make the work stand out.

Conformal Detection
Jailbreak Defense
Sequential Testing
Unseen Attack Sources
Persist-3
πŸ”Ž Similar Papers
No similar papers found.