Score
Designs and executes empirical evaluations and attack frameworks to analyze the robustness of content-moderation systems and safety classifiers, including constructing simulated or multi-turn audit trajectories, fabricated moderation traces, and progressive jailbreak strategies to elicit harmful outputs. Builds benchmarks and metrics that quantify toxicity-bypass effects and per-category vulnerabilities, measure attack success across models, and translate those findings into concrete guidance for hardening model safety constraints.
Jailbreaking attacks against large language models (LLMs) have grown increasingly sophisticated across the evolution of LLM ecosystems—from single-modality LLMs to multimodal LLMs (MLLMs) and, most recently, agentic systems—yet systematic security analysis, especially for agent architectures, remains scarce. Method: We propose the first jailbreaking threat taxonomy tailored to agent-based systems, integrating attack impact/visibility, response timing, and technical pathways. Through empirical evaluation, comprehensive literature review (2020–2025), and comparative framework analysis, we construct a unified knowledge graph of jailbreaking techniques. Contribution/Results: Our work fills a critical gap in agent-specific security analysis, reveals key limitations—including poor experimental reproducibility and lagging evaluation methodologies—and provides foundational insights for developing robust evaluation benchmarks, cross-modal defense generalization strategies, and automated red-teaming/blue-teaming platforms.
Large language models (LLMs) pose dual societal risks—unintended or malicious generation of toxic, biased, and offensive content—constituting an urgent socio-technical challenge. To address this, we conduct a systematic literature review and technical analysis to propose the first unified taxonomy for LLM safety, covering both textual and multimodal scenarios; it integrates harm categories (e.g., jailbreaking, multimodal misuse) and defense mechanisms (e.g., RLHF, prompt engineering, LLM-augmented detection). Our analysis identifies critical limitations in current evaluation methodologies—particularly regarding dynamism, cross-modal generalizability, and human-AI collaborative governance. We therefore advocate a forward-looking research agenda centered on dynamic defense strategies and human-in-the-loop governance. This work delivers a systematic theoretical framework and an extensible research paradigm for advancing LLM safety alignment.
本文审计了本地部署的大语言模型在面对越狱攻击时的防御机制,通过追踪失败原因至具体假设,并测试了六种开放权重模型。
This work addresses a structural security vulnerability in large language models (LLMs) operating in function-calling scenarios, where ambiguous trust boundaries render conventional prompt-based defenses ineffective against multi-turn jailbreaking attacks. The authors propose SMT (Simulated Moderation Traces), a black-box attack framework that progressively weakens model safety constraints by crafting multi-round dialogues masquerading as legitimate moderation processes. SMT leverages structured tool-call contexts, forged moderation frames, and feedback-driven iterative optimization to bypass safeguards across conversation turns. This approach uncovers a previously unexplored attack surface in function-calling LLMs, transcending the limitations of single-turn prompt injection paradigms. Evaluated on five leading commercial models, SMT achieves the highest average attack success rate and HarmScore with the fewest queries, underscoring the critical need for context-aware security validation mechanisms.
This study addresses the lack of systematic quantitative evaluation methods for assessing the safety robustness of large language models under jailbreak attacks. To this end, we construct an end-to-end evaluation pipeline that employs a taxonomy-driven attack selection strategy based on the AILuminate benchmark, integrating human annotation with automated calibration mechanisms. Furthermore, we introduce a "resilience gap" metric to precisely quantify the divergence in safety performance between baseline and adversarial conditions. This work establishes a reproducible comparative benchmark alongside a risk disclosure framework. Experimental results demonstrate that models exhibit an average resilience gap of 7.57%, with unsafe response rates increasing significantly from 11.08% to 18.65%, thereby revealing substantial safety vulnerabilities in current large language models.
This study addresses the lack of systematic internal analysis in current safety audits of large language models, which often fail to uncover deep-seated vulnerabilities. To bridge this gap, the work proposes an interpretability-driven activation intervention method by integrating Universal Steering with Representation Engineering for the first time. An adaptive two-stage grid search strategy is designed to optimize intervention parameters, enabling jailbreaking audits across eight prominent open-source models. Experimental results reveal that the Llama-3 series exhibits high susceptibility (with a jailbreak success rate of up to 91%), while GPT-oss-120B demonstrates robustness. Notably, Qwen and Phi models show significant performance disparities across scales. These findings validate the method’s efficacy in linking internal model representations to unsafe behaviors and underscore the dual-edged nature of interpretability techniques in security auditing.
To address the lack of systematic and quantitative evaluation for jailbreak attacks against large language models (LLMs), this paper proposes the first dedicated quantitative assessment framework for jailbreak prompts. Methodologically, it moves beyond conventional binary robustness evaluation by introducing a dual-dimensional 0–1 scoring system: coarse-grained (success in bypassing safety alignment) and fine-grained (harm severity, semantic stealthiness, etc.). The framework integrates a multi-model collaborative scoring pipeline, adversarial prompt semantic similarity analysis, and human-in-the-loop verification. Additionally, we release the first high-quality, manually curated jailbreak prompt benchmark, covering diverse attack strategies. Experiments demonstrate that our framework significantly improves detection sensitivity—identifying high-risk jailbreak prompts missed by traditional methods—and substantially enhances LLM safety risk discovery capability.
本文评估了53个大型语言模型在11个数据集上的内容安全性能,揭示了模型在不同类别有害内容上的局限性,挑战了规模即安全的假设。
本文提出EvoHarmBench,通过迭代优化生成人类可读的逃避策略,解决有害内容检测系统在线上环境下的性能差距问题。
Current safety evaluations of text-to-image models are largely confined to controlled experimental settings, failing to capture real-world risks in open ecosystems. This work presents a large-scale empirical study assessing over 200 open-source models on Hugging Face through automated safety testing, employing three representative jailbreaking attack methods. To address the overestimation of risk caused by semantic drift and visual artifacts in conventional detection, we introduce the Advanced ASR metric, which integrates semantic validity and visual plausibility criteria. Our analysis reveals, for the first time, a non-uniform degradation of model safety in open environments: while many downstream models exhibit inherent robustness even without explicit safeguards, we also identify a subset of high-risk models—including those explicitly oriented toward NSFW content and others that appear benign but are in fact unsafe. Relevant cases have been reported to the platform.
This study investigates whether existing large-model-based commercial image moderation systems are robust against evasion attacks employing simple image transformations. It presents the first systematic evaluation of three leading APIs under seven black-box image manipulations—such as color inversion and grayscale conversion—that require no gradients, surrogate models, or internal system knowledge. The findings reveal that even fixed transformations easily interpretable by humans can significantly bypass these moderation systems, with particularly pronounced vulnerabilities in multimodal content and self-harm categories. These results challenge the feasibility of relying solely on large-model APIs as standalone security boundaries and underscore their insufficiency for constructing dependable content safety mechanisms.
Current evaluations of jailbreaking attacks predominantly focus on attack success rates, which poorly reflect their actual value in improving model safety and alignment. This work proposes the first defender-centric evaluation paradigm, treating jailbreak samples as red-teaming resources and measuring attack utility through their downstream effectiveness in enhancing model robustness. To this end, we introduce the A-MESS framework, which employs an AttackSHAP score—based on Shapley values—to quantify the marginal contribution of individual attacks. By integrating black-box utility estimation with surrogate model optimization, our approach efficiently selects high-value attack subsets under query budget constraints. Experiments demonstrate that attack success rate exhibits weak correlation with defensive utility, that AttackSHAP can be accurately estimated with few queries, and that the selected subsets significantly improve model safety.