black-box adversarial attacks

Designs, builds, and analyzes methods that cause a target model to produce adversarial or unintended outputs without access to its internal parameters or gradients. This includes creating query-efficient optimization and probing strategies, automated typographic/transformational or suffix-based jailbreaks that exploit model feedback to minimize queries, techniques to transfer attacks across deployed systems, and evaluation protocols for attack success and robustness.

black-boxadversarialattacks

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.05
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$204K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Stealthy Jailbreak Attacks on Large Language Models via Benign Data Mirroring

Oct 28, 2024
HM
Honglin Mu
🏛️ Harbin Institute of Technology | Central South University | East China Normal University | Beihang University | MBZUAI | Tsinghua University

To address the vulnerability of malicious prompts to content moderation and their insufficient stealth in black-box jailbreaking attacks, this paper proposes a stealthy attack paradigm that avoids submitting detectable malicious instructions. Our method first distills benign data to construct a lightweight surrogate model approximating the target LLM’s behavior; this surrogate then guides transfer-based prompt optimization to efficiently discover jailbreaking prompts without triggering moderation. Crucially, the approach eliminates reliance on target-model feedback and reduces the risk of high-frequency malicious queries inherent in conventional black-box methods. Experiments on an AdvBench subset against GPT-3.5 Turbo achieve a 92% jailbreaking success rate, requiring only 1.5 detectable queries on average. The method attains an 80% balance score between success rate and stealth, significantly enhancing both practicality and robustness of jailbreaking attacks.

Enhancing stealth in jailbreak attacks on LLMsImproving transfer attack success rates on black-box modelsReducing detectable malicious queries during attacks

This work addresses the lack of a unified understanding of how the success rate of jailbreaking attacks on large language models systematically varies with the attacker’s computational investment. We propose the first scaling law framework for jailbreaking attacks, unifying four representative attack paradigms—optimization-based attacks, self-optimizing prompts, sampling-and-selection, and genetic optimization—under a common formulation as compute-constrained optimization processes. Evaluating these methods across multiple models and harmful objectives along a standardized FLOPs axis, we fit their success rates using saturated exponential functions. Our analysis reveals that prompt-based approaches achieve significantly higher computational efficiency and stealth, and that misleading-type harmful content is markedly easier to elicit, highlighting a strong dependence of attack efficacy on the nature of the target harmful behavior.

attack successharm typesjailbreak attacks

JailbreakLens: Visual Analysis of Jailbreak Attacks Against Large Language Models

Apr 12, 2024
YF
Yingchaojie Feng
🏛️ Zhejiang University

To address the lack of systematic analysis tools for jailbreaking attacks against large language models (LLMs), this paper introduces JailbreakLens—the first collaborative analysis framework integrating LLM-based reasoning with multidimensional visualization. It enables automated evaluation of jailbreak prompts, component-level semantic decomposition (e.g., intent, obfuscation, and trigger mechanisms), and interactive prompt refinement, supported by heatmaps, treemaps, and temporal trajectory visualizations. Its key innovations include LLM-assisted feature parsing and a human-in-the-loop verification闭环. Evaluated through case studies, technical benchmarks, and expert interviews, JailbreakLens significantly improves jailbreak pattern identification accuracy (+32.7%) and accelerates model vulnerability localization (58% reduction in average analysis time). The framework establishes a new paradigm for interpretable, reproducible LLM security assessment.

Analyzing jailbreak prompts to assess LLM security vulnerabilitiesIdentifying model weaknesses through visual and multi-level analysisStreamlining evaluation of jailbreak performance and prompt characteristics

This work addresses the security vulnerability of large language models (LLMs) to jailbreaking attacks by uncovering, from a representation engineering perspective, an intrinsic mechanism: jailbreak success or failure is not solely determined by output-layer logic but stems from a specific, semantically irrelevant yet highly detectable and controllable neural activation pattern in the latent space. The authors propose a lightweight, contrastive-query-based paradigm for identifying and intervening in this pattern—requiring only a small set of contrastive examples to reliably localize it. By selectively attenuating or amplifying its activation strength, the model’s jailbreak robustness can be significantly increased or decreased. Extensive experiments demonstrate the consistency of this mechanism across multiple mainstream open-source LLMs. The approach offers a novel, interpretable, and low-overhead pathway for enhancing LLM security, bridging representation-level analysis with practical adversarial robustness.

Exploring self-safeguarding mechanisms in LLMs.Manipulating LLM robustness against malicious inputs.Understanding LLM vulnerabilities to jailbreaking attacks.

An Optimizable Suffix Is Worth A Thousand Templates: Efficient Black-box Jailbreaking without Affirmative Phrases via LLM as Optimizer

Aug 21, 2024
WJ
Weipeng Jiang
🏛️ Xi’an Jiaotong University | Rutgers University | University of Massachusetts Amherst

Existing jailbreaking methods rely on white-box access, handcrafted templates, or inefficient search strategies, struggling to balance generality and efficiency. This paper proposes ECLIPSE—a highly efficient black-box jailbreaking framework that uniquely employs the target large language model (LLM) itself as an optimizer. It autonomously generates and iteratively refines adversarial suffixes via natural-language instructions, requiring no gradient information, predefined phrases, or human intervention. Its core innovations are: (1) an LLM-driven, self-supervised suffix optimization paradigm; and (2) a reinforcement feedback mechanism grounded in harm score estimation, enabling model introspection and rapid convergence. Evaluated on five mainstream models—including GPT-3.5-Turbo—ECLIPSE achieves a 92% average attack success rate, outperforming GCG by 2.4× in effectiveness and 83% in efficiency, while matching the performance of state-of-the-art template-based methods.

Enhancing efficiency in black-box jailbreaking without white-box accessOvercoming limitations of existing jailbreaking methods for LLMsReducing manual effort in generating harmful content templates

Latest Papers

What's happening recently
View more

Current evaluations of jailbreak attacks on large language models lack a unified standard, leading to unreliable estimates of attack success rates. This work proposes JailMeter, a novel framework that introduces, for the first time, an evidence-driven evaluation paradigm grounded in information bottleneck theory. JailMeter employs a dual-feedback optimization mechanism to filter out irrelevant noise from model responses, retaining only content pertinent to malicious intent, thereby enabling rigorous determination of whether a jailbreak has genuinely succeeded. Furthermore, the evaluator is distilled into a lightweight small language model, significantly reducing computational overhead while maintaining high reliability. Evaluated on JailMeter-Eva—a benchmark comprising 330 human-annotated samples—JailMeter achieves an assessment accuracy of 97.27%, substantially outperforming existing methods.

attack success rateevaluation frameworkjailbreak attacks

This work addresses the vulnerability of large language models (LLMs) to jailbreak attacks by proposing a novel, lightweight pre-defense mechanism. Existing pre-defense approaches suffer from high false-negative rates due to their reliance solely on user prompts, while post-hoc defenses incur substantial computational overhead. To overcome these limitations, the authors introduce a method that leverages a small language model (SLM) to generate draft responses, which—combined with the original user prompt—are fed into a security detection module. By analyzing the transferability of jailbreak attacks between LLMs and SLMs, the approach integrates speculative inference into safety verification. This strategy significantly improves detection accuracy and reduces both false negatives and computational costs, achieving an efficient and low-latency defense without compromising responsiveness.

false-negative ratejailbreak attackslarge language models

Despite the deployment of alignment and content moderation mechanisms, large language models (LLMs) remain vulnerable to black-box jailbreaking attacks. This work introduces, for the first time, a systematic application of genetic algorithms to black-box LLM jailbreaking, efficiently generating high-fitness adversarial suffixes by iteratively performing selection, mutation, and crossover operations in a discrete prompt space—without requiring access to internal model information. The proposed method successfully jailbreaks multiple mainstream commercial LLMs under realistic black-box conditions, substantially exposing the fragility of current safety safeguards and demonstrating the effectiveness and practical utility of evolution-inspired search strategies in adversarial prompt engineering.

adversarial attackblack-box settingLLM jailbreaking

This work addresses the vulnerability of generative models to jailbreaking attacks and the prohibitive cost of comprehensive evaluation. The authors propose a behavioral geometry framework that models the structural relationships among model behaviors, enabling high-accuracy prediction of cross-model jailbreak susceptibility using only a minimal set of probes. By integrating behavioral geometry modeling, vulnerability prediction, and probe optimization, the method achieves an AUPRC of 0.94 across 79 models and 100 configurations while reducing probe usage by 98%. Notably, defense strategies transferred from just three source models outperform provider-matched baselines by 2% (p = 0.03), demonstrating both efficiency and effectiveness in scalable jailbreak mitigation.

behavioral geometrydefense transfergenerative models

Hot Scholars

XC

Xiaochun Cao

Sun Yat-sen University
Computer VisionArtificial IntelligenceMultimediaMachine Learning
SL

Siyuan Liang

College of Computing and Data Science, Nanyang Technological University
Trustworthy Foundation Model
SJ

Shouling Ji

Professor, Zhejiang University & Georgia Institute of Technology
Data-driven SecurityAI SecuritySoftware ScurityPrivacy
YC

Yejin Choi

Stanford University / NVIDIA
Natural Language ProcessingDeep LearningArtificial IntelligenceCommonsense Reasoning
MB

Michael Backes

Chairman and Founding Director of the CISPA Helmholtz Center for Information Security
SecurityprivacycryptographyAI