moderation robustness analysis

Designs and executes empirical evaluations and attack frameworks to analyze the robustness of content-moderation systems and safety classifiers, including constructing simulated or multi-turn audit trajectories, fabricated moderation traces, and progressive jailbreak strategies to elicit harmful outputs. Builds benchmarks and metrics that quantify toxicity-bypass effects and per-category vulnerabilities, measure attack success across models, and translate those findings into concrete guidance for hardening model safety constraints.

moderationrobustnessanalysis

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.16
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation

Aug 07, 2025
CZ
Chi Zhang
🏛️ University of South Florida | Missouri University of Science and Technology

Large language models (LLMs) pose dual societal risks—unintended or malicious generation of toxic, biased, and offensive content—constituting an urgent socio-technical challenge. To address this, we conduct a systematic literature review and technical analysis to propose the first unified taxonomy for LLM safety, covering both textual and multimodal scenarios; it integrates harm categories (e.g., jailbreaking, multimodal misuse) and defense mechanisms (e.g., RLHF, prompt engineering, LLM-augmented detection). Our analysis identifies critical limitations in current evaluation methodologies—particularly regarding dynamism, cross-modal generalizability, and human-AI collaborative governance. We therefore advocate a forward-looking research agenda centered on dynamic defense strategies and human-in-the-loop governance. This work delivers a systematic theoretical framework and an extensible research paradigm for advancing LLM safety alignment.

Improve safety via RLHF prompt engineering alignmentLLMs generate harmful toxic offensive biased contentNeed defenses against adversarial jailbreaking attacks

Must-Read Papers

Most classic and influential ideas
View more

This work addresses a structural security vulnerability in large language models (LLMs) operating in function-calling scenarios, where ambiguous trust boundaries render conventional prompt-based defenses ineffective against multi-turn jailbreaking attacks. The authors propose SMT (Simulated Moderation Traces), a black-box attack framework that progressively weakens model safety constraints by crafting multi-round dialogues masquerading as legitimate moderation processes. SMT leverages structured tool-call contexts, forged moderation frames, and feedback-driven iterative optimization to bypass safeguards across conversation turns. This approach uncovers a previously unexplored attack surface in function-calling LLMs, transcending the limitations of single-turn prompt injection paradigms. Evaluated on five leading commercial models, SMT achieves the highest average attack success rate and HarmScore with the fewest queries, underscoring the critical need for context-aware security validation mechanisms.

function-calling LLMsjailbreak attacksmulti-turn execution

This study addresses the lack of systematic quantitative evaluation methods for assessing the safety robustness of large language models under jailbreak attacks. To this end, we construct an end-to-end evaluation pipeline that employs a taxonomy-driven attack selection strategy based on the AILuminate benchmark, integrating human annotation with automated calibration mechanisms. Furthermore, we introduce a "resilience gap" metric to precisely quantify the divergence in safety performance between baseline and adversarial conditions. This work establishes a reproducible comparative benchmark alongside a risk disclosure framework. Experimental results demonstrate that models exhibit an average resilience gap of 7.57%, with unsafe response rates increasing significantly from 11.08% to 18.65%, thereby revealing substantial safety vulnerabilities in current large language models.

AI safety evaluationJailbreak attacksLarge language models

This study addresses the lack of systematic internal analysis in current safety audits of large language models, which often fail to uncover deep-seated vulnerabilities. To bridge this gap, the work proposes an interpretability-driven activation intervention method by integrating Universal Steering with Representation Engineering for the first time. An adaptive two-stage grid search strategy is designed to optimize intervention parameters, enabling jailbreaking audits across eight prominent open-source models. Experimental results reveal that the Llama-3 series exhibits high susceptibility (with a jailbreak success rate of up to 91%), while GPT-oss-120B demonstrates robustness. Notably, Qwen and Phi models show significant performance disparities across scales. These findings validate the method’s efficacy in linking internal model representations to unsafe behaviors and underscore the dual-edged nature of interpretability techniques in security auditing.

interpretabilityjailbreakinglarge language models

AttackEval: How to Evaluate the Effectiveness of Jailbreak Attacking on Large Language Models

Jan 17, 2024
DS
Dong Shu
🏛️ Northwestern University | Rutgers University | University of Liverpool

To address the lack of systematic and quantitative evaluation for jailbreak attacks against large language models (LLMs), this paper proposes the first dedicated quantitative assessment framework for jailbreak prompts. Methodologically, it moves beyond conventional binary robustness evaluation by introducing a dual-dimensional 0–1 scoring system: coarse-grained (success in bypassing safety alignment) and fine-grained (harm severity, semantic stealthiness, etc.). The framework integrates a multi-model collaborative scoring pipeline, adversarial prompt semantic similarity analysis, and human-in-the-loop verification. Additionally, we release the first high-quality, manually curated jailbreak prompt benchmark, covering diverse attack strategies. Experiments demonstrate that our framework significantly improves detection sensitivity—identifying high-risk jailbreak prompts missed by traditional methods—and substantially enhances LLM safety risk discovery capability.

Create a ground truth dataset for jailbreak prompt evaluation.Develop frameworks for coarse-grained and fine-grained attack assessments.Evaluate jailbreak attack effectiveness on large language models.

Latest Papers

What's happening recently
View more

Current safety evaluations of text-to-image models are largely confined to controlled experimental settings, failing to capture real-world risks in open ecosystems. This work presents a large-scale empirical study assessing over 200 open-source models on Hugging Face through automated safety testing, employing three representative jailbreaking attack methods. To address the overestimation of risk caused by semantic drift and visual artifacts in conventional detection, we introduce the Advanced ASR metric, which integrates semantic validity and visual plausibility criteria. Our analysis reveals, for the first time, a non-uniform degradation of model safety in open environments: while many downstream models exhibit inherent robustness even without explicit safeguards, we also identify a subset of high-risk models—including those explicitly oriented toward NSFW content and others that appear benign but are in fact unsafe. Relevant cases have been reported to the platform.

in-the-wild safetyjailbreakmodel release practices

This study investigates whether existing large-model-based commercial image moderation systems are robust against evasion attacks employing simple image transformations. It presents the first systematic evaluation of three leading APIs under seven black-box image manipulations—such as color inversion and grayscale conversion—that require no gradients, surrogate models, or internal system knowledge. The findings reveal that even fixed transformations easily interpretable by humans can significantly bypass these moderation systems, with particularly pronounced vulnerabilities in multimodal content and self-harm categories. These results challenge the feasibility of relying solely on large-model APIs as standalone security boundaries and underscore their insufficiency for constructing dependable content safety mechanisms.

AI safetycontent moderationfoundation models

Current evaluations of jailbreaking attacks predominantly focus on attack success rates, which poorly reflect their actual value in improving model safety and alignment. This work proposes the first defender-centric evaluation paradigm, treating jailbreak samples as red-teaming resources and measuring attack utility through their downstream effectiveness in enhancing model robustness. To this end, we introduce the A-MESS framework, which employs an AttackSHAP score—based on Shapley values—to quantify the marginal contribution of individual attacks. By integrating black-box utility estimation with surrogate model optimization, our approach efficiently selects high-value attack subsets under query budget constraints. Experiments demonstrate that attack success rate exhibits weak correlation with defensive utility, that AttackSHAP can be accurately estimated with few queries, and that the selected subsets significantly improve model safety.

attack utilitydefender-centric evaluationjailbreak attacks

Hot Scholars

LL

Luca Luceri

Research Assistant Professor @University of Southern California - Information Sciences Institute
Computational Social ScienceNetwork ScienceMachine LearningSocial Media Manipulation
GN

Gianluca Nogara

SUPSI
Computational Social ScienceData ScienceOnline Social Networks
SG

Silvia Giordano

Prof. Networking, SUPSI
networkingperformancespervasive and mobile computing
FM

Filippo Menczer

Luddy Distinguished Professor of Informatics and Computer Science, Indiana University
MisinformationWeb ScienceNetwork ScienceComputational Social Science
HW

Hanlin Wu

Tsinghua University
Generative ModelsAI for Science