Score
Designs and executes experiments and analyses to measure how well security or privacy countermeasures reduce attacks or information leakage, quantify absolute and relative mitigation effectiveness, and detect residual vulnerabilities. Compares alternative countermeasure designs by quantifying tradeoffs such as performance, resource cost, and security gains to support selection and improvement.
This study addresses the significant abstraction gap between security-by-design specifications—typically expressed in domain-specific languages (DSLs)—and code-level analyzers, which impedes the traceability of design intent to implementation vulnerabilities. It presents the first large-scale empirical investigation, examining 559 security checks across 36 analyzers and 66 security design DSLs. The authors introduce SecLan, a unified model that captures shared security concepts between these two layers, and validate its structure through expert evaluation involving 22 practitioners and qualitative interviews with 9 additional experts. The findings reveal a pronounced mismatch between security concepts at the design and implementation levels, with existing analyzer checks often relying on overly broad vulnerability descriptions, leading to ambiguous mappings. This work provides both an empirical foundation and a modeling framework to bridge the gap between security design and implementation.
This study addresses the widespread issue in AI red-teaming evaluations where comparisons of attack success rates (ASR) often rely on invalid or incomparable measurements, leading to erroneous conclusions about system security or attack efficacy. For the first time, the paper introduces measurement validity theory from social sciences into the AI red-teaming domain, systematically analyzing—through inferential statistics and illustrative case studies such as jailbreaking attacks—the conditions under which ASR comparisons are meaningful. The work establishes clear prerequisites for valid ASR comparisons, identifies and categorizes common patterns of invalid comparison, and thereby provides a rigorous theoretical foundation that significantly enhances the scientific rigor and comparability of AI safety evaluations.
This study investigates the effectiveness of code obfuscation techniques in impeding attackers’ comprehension of malicious logic and examines whether quantitative code complexity metrics can predict their impact on attack success rates and time-to-compromise. Through a controlled user study, it systematically evaluates, for the first time, the cumulative defensive efficacy of multiple obfuscation techniques—including control-flow flattening, string encryption, and virtualization—using both quantitative measures (e.g., comprehension time, success rate) and qualitative feedback (e.g., cognitive load assessments). It establishes an empirically validated link between objective complexity metrics (e.g., cyclomatic complexity, AST depth) and subjective attack difficulty. Results show that multi-layered obfuscation significantly increases attacker comprehension time (average +217%) and that certain metrics—particularly control-flow entropy—effectively predict attack failure probability. The work provides reproducible, evidence-based guidance for obfuscation strategy selection and software protection evaluation.
Current cyber-physical systems (CPS) in vehicular environments lack quantitative, experimentally grounded methods for assessing network resilience. Method: This study constructs an experimental testbed replicating real-world truck operational conditions and conducts multiple rounds of malware injection attacks, simultaneously collecting network- and physical-layer data on resistance and recovery behaviors. Contribution/Results: We introduce the novel concept of “bonware” to holistically characterize both cybersecurity defense capability and physical resilience, formalized via an analytically tractable mathematical model. We further define and extract experimentally identifiable, quantitative resilience metrics—termed elastic features—for the first time. Sensitivity analysis confirms these metrics exhibit significant discriminability with respect to attack intensity, defensive strategies, and physical redundancy. This work bridges a critical gap by advancing vehicular CPS resilience from qualitative description to quantifiable, comparable, and optimizable measurement.
This work addresses the limitations of existing adversarial simulation tools, which rely on agent-based instrumentation of target systems, often leaving anomalous artifacts and failing to faithfully replicate human attacker behavior—particularly in critical phases of the cyber kill chain such as initial access and interactive operations. To overcome these shortcomings, the authors propose and implement an open-source attack scripting language coupled with an agentless execution engine that closely emulates real-world attacker tactics. This approach enables high-fidelity, interactive simulation of complete kill chain stages, including initial access, privilege escalation, and lateral movement. Experimental results demonstrate that system logs generated by this method exhibit significantly greater behavioral similarity to those produced by actual human-driven attacks, thereby enhancing the realism and effectiveness of security testing and intrusion detection research.
This work addresses a critical limitation in existing end-to-end black-box evaluations of large language models (LLMs) for automated exploitation, where errors in the reconnaissance phase obscure the true exploit capabilities of LLMs. To resolve this, the authors propose a two-stage decoupled evaluation framework that isolates reconnaissance and exploitation performance by injecting real-world vulnerability contexts and applying knowledge-driven ablation. Evaluated across 70 high-fidelity web vulnerability environments, the framework enables the first independent quantification of these two capabilities. Comparative analysis across 50 representative vulnerabilities reveals that, given accurate contextual information, LLMs achieve up to 90% exploit success rates, whereas autonomous reconnaissance yields only ~50% recall. Furthermore, multi-agent, monolithic, and graph-driven architectures exhibit distinct strengths and limitations across vulnerability types involving long-sequence interactions, short-chain injections, and cross-session access control, thereby delineating their respective capability boundaries.
This study addresses a critical gap in the evaluation of security agents, which has traditionally emphasized success rates while overlooking the operational costs associated with reasoning steps, tool invocations, and telemetry queries—factors essential to real-world economic efficiency. To bridge this gap, the authors propose the first cost-aware evaluation framework tailored for security operations, integrating Cybench offensive tasks and Splunk BOTS v1 defensive scenarios under fixed resource budgets. The framework enables fine-grained quantification of computational and query-related overheads. Findings reveal that open-source large language models can match or surpass proprietary systems in offensive tasks at lower cost, whereas defensive performance hinges critically on efficient tool utilization rather than raw computational power. These insights, disseminated via an interactive website, uncover fundamental differences in scaling dynamics between red- and blue-team operations.
Real-world Security Operations Center (SOC) data is rarely accessible for research due to privacy constraints, leading existing studies to rely on synthetic or outdated datasets. This work proposes a high-fidelity anonymization method that extracts and structures SIEM logs from a financial-sector SOC, preserving temporal ordering and entity consistency while enforcing strict privacy guarantees—thereby establishing the first quantifiable privacy-utility trade-off boundary. Leveraging this approach, we construct 37 HIKARI evaluation challenges and develop a deterministic validator alongside a large language model (LLM) behavioral compliance detection mechanism. In experiments involving 200 SOCpilot incidents, our framework uncovered LLM non-compliant actions undetected by human baselines, enabling reproducible and verifiable evaluation of autonomous defense systems.
This study addresses the inherent trade-off between security and fidelity in large language models when defending against indirect prompt injection attacks. Such defenses often suppress untrusted input, inadvertently degrading performance on tasks requiring high output faithfulness, such as translation and text editing. The work presents SecFid, the first evaluation benchmark capable of distinguishing whether a model executes, retains, or ignores injected content. Through large-scale experiments encompassing 1,168 samples across 48 configurations, decision-theoretic analysis, and multi-model comparisons, the authors quantitatively demonstrate the tension between these objectives: the highest-fidelity model achieves 96.5% fidelity but only 47.8% security, whereas the most secure model attains 99.3% security at the cost of fidelity dropping to 71.0%–73.9%, confirming that simultaneously optimizing both dimensions remains fundamentally challenging.