mitigation effectiveness testing

Designs and run experiments, benchmarks, and metrics to quantify how well a proposed mitigation or defensive mechanism reduces unwanted or malicious behavior and to compare alternative defenses. Builds test harnesses, threat models, datasets or workloads, and evaluation procedures to measure reduction in misuse and to analyze residual or covert vulnerabilities under realistic and adversarial conditions.

mitigationeffectivenesstesting

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.77
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$192K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Benchmarking Misuse Mitigation Against Covert Adversaries

Jun 06, 2025
DB
Davis Brown
🏛️ University of Pennsylvania | Carnegie Mellon University

Existing language model safety evaluations primarily target single-turn explicit attacks, failing to address emerging threats where adversaries decompose hazardous tasks into numerous seemingly benign, independent queries and stealthily achieve malicious objectives via state accumulation across interactions. Method: We formally define the “decomposition-based covert attack” paradigm and propose BSD—a state-aware, multi-turn interaction modeling and automated attack data generation framework—to enable systematic evaluation of stateful defense capabilities. Contribution/Results: We introduce the first benchmark framework supporting state-aware safety assessment, releasing two high-difficulty datasets on which leading closed-source models consistently refuse and open-source models largely fail. We further propose a cross-model security response consistency metric. Experiments demonstrate that decomposition-based attacks successfully compromise mainstream models, while state-preserving defenses significantly enhance resilience against covert misuse—establishing a quantifiable, reproducible evaluation infrastructure for safety alignment.

Benchmarking model resilience to hidden multi-query exploitation strategiesDetecting covert misuse of language models through fragmented queriesEvaluating defenses against stateful adversarial attacks on AI systems

This work addresses the limitations of existing vulnerability detection benchmarks, which are predominantly confined to the function level and thus fail to capture cross-procedural vulnerabilities prevalent in real-world scenarios, while manually curated repository-scale datasets suffer from limited scalability. To overcome these challenges, the paper proposes an automated benchmark generation framework that injects realistic vulnerabilities into genuine code repositories and synthesizes reproducible proofs of vulnerability (PoVs), thereby constructing a large-scale, precisely labeled repository-level vulnerability dataset. This framework represents the first scalable approach to automatically generating repository-level vulnerability benchmarks. Furthermore, it introduces an adversarial co-evolution mechanism, enabling dynamic interaction between vulnerability injection and detection agents under realistic constraints, which significantly enhances the robustness and practical utility of vulnerability detection models.

automated dataset generationrealistic vulnerabilityrepository-level benchmark

This study addresses the lack of objective validation criteria in existing threat modeling approaches, which often rely on expert judgment and are thus prone to omissions or inconsistencies. To overcome this limitation, the authors propose a quantifiable and reproducible evaluation methodology based on benchmark applications with known vulnerabilities—specifically AzureGoat and VulnBank. Using only architectural diagrams, data flow diagrams, and their textual descriptions as input, the approach evaluates the vulnerability coverage of ThreMoLIA, an LLM-assisted threat modeling system, against Microsoft Threat Modeling Tool. Experimental results demonstrate that ThreMoLIA achieves consistently higher vulnerability coverage across both benchmark applications. This work represents the first effort to employ real-world vulnerable applications as a validation benchmark for threat modeling, effectively mitigating the shortcomings inherent in traditional expert-based assessments.

completeness assessmentsecurity evaluationtest applications

AI safety evaluation lacks consensus standards, limiting its utility for governance and policy decisions. This paper introduces the first practical AI safety evaluation framework, systematically integrating threat modeling, assessment design, and validity validation. It formally defines three essential criteria for “useful” evaluations—risk alignment, reproducibility, and scalability—along with associated quantitative parameters. Innovatively distinguishing formal metrics from real-world risk coverage, the framework establishes an evolutionary paradigm—from isolated tests to modular, composable evaluation suites. It synergistically integrates red-teaming, evaluation validity analysis, and cybersecurity best practices to jointly optimize reliability, construct validity, and operational feasibility. The resulting safety evaluation guidelines have been adopted by industry stakeholders and policymaking bodies, demonstrably enhancing the interpretability of evaluation outcomes and their actionable support for risk-informed decision-making.

Connecting threat modeling to evaluation design effectivelyDefining criteria for high-quality AI safety evaluationsDeveloping comprehensive evaluation suites for AI systems

This work addresses the limitations of existing adversarial simulation tools, which rely on agent-based instrumentation of target systems, often leaving anomalous artifacts and failing to faithfully replicate human attacker behavior—particularly in critical phases of the cyber kill chain such as initial access and interactive operations. To overcome these shortcomings, the authors propose and implement an open-source attack scripting language coupled with an agentless execution engine that closely emulates real-world attacker tactics. This approach enables high-fidelity, interactive simulation of complete kill chain stages, including initial access, privilege escalation, and lateral movement. Experimental results demonstrate that system logs generated by this method exhibit significantly greater behavioral similarity to those produced by actual human-driven attacks, thereby enhancing the realism and effectiveness of security testing and intrusion detection research.

adversary emulationattack automationcyber attack scenarios

Latest Papers

What's happening recently
View more

This study addresses the lack of a systematic evaluation framework for AI-driven Windows malware detectors, which hinders effective model selection in real-world deployments. To bridge this gap, the authors propose EXE-Bench, a comprehensive benchmark that unifies performance, temporal robustness, adversarial robustness, and computational overhead into a single scoring system. By integrating multidimensional metrics, temporal evolution analysis, content-injection adversarial attacks, and inference resource measurements, EXE-Bench enables fair and holistic model comparison. The evaluation reveals that feature engineering methods grounded in domain knowledge significantly outperform most deep learning models under long-term deployment and adversarial conditions, whereas the latter exhibit superior performance only in initial stages. These findings highlight the limitations of relying solely on post-deployment evaluation and reaffirm the enduring value of expert-driven feature engineering in malware detection.

adversarial robustnessAI-based detectorscomputational overhead

This study addresses the limitation of existing LLM agent safety evaluations that treat decomposition attacks and prompt injection in isolation, thereby failing to precisely localize the moment harm occurs. To bridge this gap, this work proposes a unified abuse monitoring benchmark and an action-based monitoring paradigm that formalizes the harm window by tracing externalized actions, alongside interval-based metrics for fine-grained temporal localization assessment. Validated on approximately 6,200 synthetic dialogues, the action-based monitor achieves AUCs of 0.95 and 0.99 against both attack types, whereas content-based monitors fail under injection scenarios. This research is the first to incorporate multi-source attacks into a unified modeling framework, revealing that conventional metrics significantly overestimate harm localization capabilities.

decomposition attacksharm localizationLLM agents

This study addresses the challenge of effectively evaluating and enhancing the capability of large language models (LLMs) to generate high-quality proof-of-concept (PoC) exploit code from CVE vulnerability descriptions. The authors propose a data-driven paradigm that constructs a high-quality exploit dataset through multi-stage data curation and introduces a scalable evaluation framework based on LLM-as-a-judge with fine-grained scoring criteria. They systematically assess the zero-shot performance of 17 LLMs and investigate the impact of instruction tuning and test-time rejection strategies. Experimental results demonstrate that an instruction-tuned 8B open-source model achieves over a 42.5% improvement in PoC generation quality, and when combined with a simple rejection mechanism, its performance rivals that of certain closed-source models, underscoring the critical role of data quality and thoughtful evaluation design in cybersecurity tasks.

CVE-conditionedcybersecuritydata-centric

This study addresses the lack of systematic evaluation regarding security mechanisms in coding agent frameworks and their runtime implications. We propose the first comprehensive empirical benchmark, establishing a taxonomy of ten security mechanisms. By integrating LLM-as-a-judge evaluation, deterministic oracles, and network isolation techniques, we quantify the utility and attack resilience of nine mechanisms across six mainstream frameworks using 400 cases spanning 23 tasks. Our findings reveal that automatic approval escalates the attack success rate from 29.2% to 95.6%, demonstrating that single-layer defenses are readily circumvented. Furthermore, this work provides verifiable security configuration recommendations that effectively balance legitimate task execution requirements with robust safety assurance.

AI SecurityAttack SurfacesBenchmark

Hot Scholars

KY

Kwok-Yan Lam

Nanyang Technological University
CybersecurityPrivacy-Preserving technologiesDigital TrustDistributing systems
MY

Min Yang

Bytedance
Vision Language ModelComputer VisionVideo Understanding
MC

Mauro Conti

IEEE Fellow - Prof.@University of Padua - Wallenberg WASP Guest.Prof.@Örebro U.- Affiliate Prof.@UW
SecurityPrivacy
XH

Xinlei He

Assistant Professor, HKUST(GZ)
Trustworthy Machine LearningSecurityPrivacy
TZ

Tianwei Zhang

Nanyang Technological University
Computer System Security