Score
Designs and carries out systematic analyses and evaluation frameworks that characterize and compare security posture across multiple orthogonal axes, including building and analyzing security models and performing comparative sandbox benchmarking. Produces per-product security portraits and threat-model qualification matrices, derives class-level architectural patterns, performs cross-axis security comparisons, and identifies measurement and coverage gaps.
Traditional text- and numeric-based analysis methods exhibit insufficient detection efficacy for complex software systems under dynamic threat conditions. To address this, we conduct a systematic literature review (SLR) covering over 60 key studies and propose the first visualization taxonomy tailored to software security—categorizing techniques into graph-based, symbolic, matrix-based, and metaphor-based approaches, while explicitly distinguishing development-time and runtime security application scenarios. We introduce “adaptive security visualization” as a novel research direction, emphasizing the need for real-time updates, context awareness, and human–machine collaboration under evolving threats. Our findings clarify the technical evolution trajectory and identify critical pathways to enhance threat identification accuracy and accelerate response decision-making. The work establishes a foundational theoretical framework and methodological guidance for both research and practice in security visualization.
This work addresses the challenge of detecting multi-stage attacks that traverse trust boundaries in cloud deployments—threats often missed by conventional security tools due to their inability to model holistic system architecture and runtime behavioral deviations. The authors propose a novel approach that integrates static configuration analysis with runtime network flow observation to automatically construct a platform-agnostic architectural abstraction reflecting the system’s true state, including components, domains, interfaces, policies, and data flows. Building upon this representation, the method enables continuous, architecture-level threat modeling. It is the first to support automated architecture inference and threat detection across bare-metal, Kubernetes, and cloud environments. Evaluated on supply chain systems incorporating machine learning (ML) components, the approach successfully identified all 17 classes of injection threats—including ML-specific threats—substantially outperforming existing tools, which cover only 6–47% of these threats and fail entirely to detect ML-related ones.
Scientific research cyberinfrastructure (CI) faces unique challenges—including high collaboration requirements, component heterogeneity, and the absence of adaptable security assessment frameworks. To address these, we propose a mission-centric security posture assessment method: first, top-down identification of critical assets and unacceptable losses; second, construction of a security knowledge graph integrating system components, dependencies, and threat behaviors; and third, integration with directed attack graphs to quantify multi-hop attack paths from entry points to critical assets—enabling visualization of attacker-defender relationships and identification of security blind spots. Unlike conventional generic standards, our approach is the first to deeply couple mission-driven assessment, knowledge graphs, and attack graphs. It supports risk prioritization and generation of actionable defensive strategies, significantly enhancing the precision and effectiveness of CI security defense.
Modern web browsers have evolved into critical business platforms, yet their client-side security posture lacks systematic assessment. This paper introduces the first purely frontend, in-browser security assessment framework, implemented in JavaScript and WebAssembly. It integrates over 120 fine-grained checks covering core mechanisms—including the Same-Origin Policy, Content Security Policy (CSP), sandboxing, and XSS protections—and uniquely incorporates previously unobservable OS- and network-layer dimensions such as WeakRef interference, SharedArrayBuffer availability, and internal network reachability. Leveraging dynamic policy injection and coordinated multi-API observation (e.g., Permissions, Crypto, and Reporting APIs), our empirical evaluation across enterprise environments reveals widespread policy degradation: CSP bypass rates reach 63% on legacy browsers, and SSL certificate validation is omitted in 41% of cases. The framework establishes a practical, evidence-based diagnostic paradigm for precise security hardening.
Current safety risk assessments predominantly rely on manual processes, rendering them susceptible to subjectivity, inconsistency, human error, and deficiencies in reproducibility and transparency. To address these limitations, this study presents the first independent replication of a peer-reviewed safety assessment case and introduces a Model-Based Systems Engineering–Computer-Aided Design (MBSE-CAD) methodology that embeds safety analysis directly into the system modeling workflow. This enables structured reconstruction and automated verification of the assessment process. Empirical evaluation—comparing conventional manual approaches with MBSE-CAD tools on real-world industrial cases—demonstrates significant improvements in analytical consistency, traceability, and reproducibility, while substantially reducing human bias. The work establishes a rigorous, scalable, and sustainable framework for safety engineering practice, bridging the gap between industrial application and academic research.
This study addresses the lack of objective validation criteria in existing threat modeling approaches, which often rely on expert judgment and are thus prone to omissions or inconsistencies. To overcome this limitation, the authors propose a quantifiable and reproducible evaluation methodology based on benchmark applications with known vulnerabilities—specifically AzureGoat and VulnBank. Using only architectural diagrams, data flow diagrams, and their textual descriptions as input, the approach evaluates the vulnerability coverage of ThreMoLIA, an LLM-assisted threat modeling system, against Microsoft Threat Modeling Tool. Experimental results demonstrate that ThreMoLIA achieves consistently higher vulnerability coverage across both benchmark applications. This work represents the first effort to employ real-world vulnerable applications as a validation benchmark for threat modeling, effectively mitigating the shortcomings inherent in traditional expert-based assessments.
This work addresses the lack of systematic evaluation benchmarks for large language models (LLMs) in security audit log investigation tasks by introducing AuditBench, the first audit log benchmark specifically designed for attack investigation. AuditBench encompasses over 50 real-world scenarios across Linux and Windows systems and focuses on four core tasks: alert classification, persistence mechanism identification, among others. Through multidimensional experiments, the study systematically evaluates the impact of model scale, log representation, prompt design, and fine-tuning strategies on performance and error patterns, while also analyzing the quality of LLM-generated explanations. The findings reveal the capability boundaries and characteristic failure modes of various models across different investigative tasks, providing empirical foundations for deploying and optimizing LLMs in security operations.
This study addresses the lack of systematic, engine-level security comparisons among existing AI code sandboxing solutions, which hinders reliable assessment of their isolation capabilities and associated risks. The work proposes a novel cross-dimensional analytical framework that evaluates five mainstream sandbox products across six criteria: host attack surface, information leakage potential, stackability of defense-in-depth mechanisms, historical CVE records, patching cadence, and upstream fuzz testing maturity. Through comprehensive attack surface mapping, information flow analysis, CVE mining, patch latency tracking, and fuzzing ecosystem evaluation, the research uncovers clear security boundaries between distinct engine architectures and significant disparities in security practices among functionally similar products. Notably, patch delays range from zero to over 471 days, and the optimal security configuration—combining microVM isolation with continuous public fuzz testing—remains unadopted. To guide practical deployment, the authors introduce a threat-model-adaptive selection matrix as a nuanced alternative to simplistic rankings.
This work addresses the current lack of a systematic understanding of the capabilities of AI sandboxes in ensuring safety, security, and regulatory compliance, particularly within physical AI and cyber-physical systems. It proposes the first unified, assurance-oriented framework for AI sandboxes, introducing a formal boundary definition, a comprehensive sandbox taxonomy, a threat model targeting the assurance mechanisms themselves, and a quantifiable evaluation methodology spanning six dimensions—including fidelity and controllability. Through formal modeling, threat analysis, and multi-case validation, the study clarifies what aspects of AI behavior can be effectively tested in sandboxes, which risk categories can be meaningfully controlled, and what forms of evidence such environments can generate to support safety and compliance claims, thereby establishing foundational tools for trustworthy AI verification.
This work addresses the absence of a public, unified, and queryable enterprise security environment for end-to-end evaluation of the trustworthiness and effectiveness of autonomous cyber defense agents. It presents the first open benchmarking framework tailored for autonomous enterprise defense, leveraging frozen snapshots of enterprise security states combined with dual-modality interfaces—text-to-SQL and native APIs—to enable reproducible and auditable assessments of agents’ investigative capabilities across multiple dimensions, including identity, cloud, and data security. The framework integrates relational snapshots, multi-vendor APIs, synthetic organizational datasets, and a multidimensional scoring mechanism. It has been successfully instantiated with two identity security packages and synthetic environments at varying scales, demonstrating its efficacy in the evaluation phase and laying the groundwork for future extension into remediation.