Score
Designing and applying concrete measures and assessments of user- and system-level outcomes to demonstrate impact and maturity, quantify how verification or controls affect reliance and decision-making, and evaluate perceptions of effectiveness under security/privacy constraints.
This work addresses the current lack of a systematic understanding of the capabilities of AI sandboxes in ensuring safety, security, and regulatory compliance, particularly within physical AI and cyber-physical systems. It proposes the first unified, assurance-oriented framework for AI sandboxes, introducing a formal boundary definition, a comprehensive sandbox taxonomy, a threat model targeting the assurance mechanisms themselves, and a quantifiable evaluation methodology spanning six dimensions—including fidelity and controllability. Through formal modeling, threat analysis, and multi-case validation, the study clarifies what aspects of AI behavior can be effectively tested in sandboxes, which risk categories can be meaningfully controlled, and what forms of evidence such environments can generate to support safety and compliance claims, thereby establishing foundational tools for trustworthy AI verification.
Traditional compliance assessments rely on point-in-time audits and self-attestation, which struggle to enable continuous, cross-organizational, and traceable verification of security controls in multi-vendor environments. This work proposes a permissioned blockchain-based Third-Party Risk Assessment (TPRA) framework that transforms static compliance into a dynamic, repeatable, and verifiable continuous governance mechanism through smart contract–automated evaluation workflows, multi-party governance protocols, and longitudinal state tracking. The study contributes an actionable TPRA architecture, along with complementary compliance maturity metrics and a qualitative model, enabling quantification and long-term validation of security control implementation maturity across organizational boundaries and time periods.
AI systems deployed in real-world settings face significant safety and efficacy risks due to misalignment among robustness, reliability, and accountability—key dimensions of trustworthy AI. Method: This study proposes the first theoretical framework that explicitly embeds *accountability* as a core dimension in AI evaluation. Through conceptual evolution analysis, systematic literature review, and multi-source empirical case studies, it develops a tripartite, synergistic assessment model integrating *robustness*, *reliability*, and *accountability*. The framework innovatively unifies governance-by-design, dynamic testing, and responsibility traceability into a cross-layer analytical paradigm. Contribution/Results: It identifies six critical technical and institutional challenges and five novel categories of testing requirements, and delineates a co-evolutionary pathway for technical capabilities and regulatory infrastructure. The work provides an actionable theoretical foundation and implementation roadmap for AI standardization, regulatory practice, and liability attribution.
Enterprises struggle to quantitatively assess the effectiveness of Zero Trust Architecture (ZTA) implementations and lack a structured, stage-wise evolution roadmap. Method: This study proposes a four-level, data-driven Zero Trust Maturity Model (ZTMM), the first to define a quantifiable set of technical controls—including identity verification, micro-segmentation, end-to-end encryption, and automated policy orchestration—grounded in industry best practices and empirically validated. Maturity is stratified into Initial, Developing, Mature, and Optimized levels. Contribution/Results: The model enables organizations to precisely diagnose their current posture, design phased transformation roadmaps, and measure improvements in security posture. Empirical validation demonstrates that adoption significantly enhances organizational resilience against advanced persistent threats (APTs). The ZTMM provides a reusable, standardized assessment framework and implementation guidance for operationalizing ZTA.
This study addresses the underexplored tension between institutional expectations and lived experience among CMMC assessors operating in non-consultative roles. Drawing on role conflict theory, it employs interpretative phenomenological analysis (IPA) to conduct semi-structured interviews with CMMC-certified assessors, systematically uncovering their subjective experiences and logics of duty fulfillment in this mode. Findings reveal that assessors navigate role conflicts through strategies centered on technical competence, procedural discipline, and boundary management. These insights not only extend theoretical understandings of professional credibility construction in cybersecurity compliance contexts but also offer empirical grounding for establishing interactional norms and boundary-setting practices within CMMC implementation frameworks.
This study addresses the limitations of existing large language model (LLM) vulnerability detection benchmarks, which rely on a single metric and fail to accommodate the diverse evaluation needs of different security stakeholders. To bridge this gap, we propose SecLens-R—the first role-oriented, multidimensional evaluation framework—defining five role-specific weighting schemes across 35 dimensions grouped into seven categories. We evaluate 12 state-of-the-art LLMs on 406 tasks spanning 10 programming languages and 8 OWASP vulnerability types, using both Code-in-Prompt and Tool-Use paradigms. Results reveal significant performance disparities across roles, with score differences up to 31 points for the same model (e.g., Qwen3-Coder: 76.3 for an Engineering Lead vs. 45.2 for a CISO), underscoring the necessity of contextualized, multi-objective assessment and advancing vulnerability detection from uniform standards toward role-driven decision-making paradigms.
This study addresses the complex assurance challenges confronting AI-enabled Cyber-Physical Systems (AI-CPS) across perception, computation, control, human factors, and governance dimensions, noting that mere compliance with ISO/IEC 42001 fails to reveal architectural impacts or practical maturity. The authors propose CEDAR-42001, a two-stage method that uniquely maps compliance audit evidence onto a seven-layer AI-CPS architecture and governance hierarchy. By integrating a five-dimensional maturity profile, constraint identification, and rule-driven reasoning, the approach generates a traceable, architecture-aware assurance posture. Applied to an autonomous vehicle fleet case, it revealed that while 89.9% of audit items were compliant, only 34.3% met a high-assurance baseline. The method successfully reconstructed the 2023 Cruise incident, precisely identifying cross-layer deficiencies and recommending targeted mitigations to inform decision-making from strategic to operational levels.
This study addresses the significant abstraction gap between security-by-design specifications—typically expressed in domain-specific languages (DSLs)—and code-level analyzers, which impedes the traceability of design intent to implementation vulnerabilities. It presents the first large-scale empirical investigation, examining 559 security checks across 36 analyzers and 66 security design DSLs. The authors introduce SecLan, a unified model that captures shared security concepts between these two layers, and validate its structure through expert evaluation involving 22 practitioners and qualitative interviews with 9 additional experts. The findings reveal a pronounced mismatch between security concepts at the design and implementation levels, with existing analyzer checks often relying on overly broad vulnerability descriptions, leading to ambiguous mappings. This work provides both an empirical foundation and a modeling framework to bridge the gap between security design and implementation.
This work addresses the limitations of existing AI trustworthiness assessment approaches, which are either too abstract to support full lifecycle monitoring or rely on single metrics insufficient for governance needs. The paper proposes a lightweight, auditable framework for dynamic trustworthiness management that integrates formal modeling with governance processes. By employing context-sensitive trustworthiness dimension protocols and interpretable rule learning based on decision trees, the framework enables end-to-end monitoring and documentation of AI systems—from design and deployment through re-evaluation. Novel diagnostic tools, including hierarchical transitions, margin-of-boundary analysis, and profile drift detection, are introduced alongside clearly accountable human-in-the-loop checkpoints. Experiments on synthetic AI lifecycle trajectories demonstrate the framework’s effectiveness in detecting performance degradation, abrupt perturbations, and impacts of system updates, thereby establishing a transparent, traceable, and contestable evidentiary basis for AI governance.
This study addresses the limitations of current AI system evaluations, which often suffer from inconsistent methodologies and metrics that yield incomparable results and poor alignment with real-world contexts and human needs. To bridge this gap, the authors propose a reproducible three-stage scenario generation pipeline that integrates human-centered design, operational feasibility, and methodological transparency. The approach begins by eliciting authentic AI use cases from domain experts via structured use case worksheets, then leverages large language model prompt engineering combined with iterative human review to transform these into human-oriented evaluation scenarios. A validation rubric is developed to assess scenario quality. Applied in the financial services sector, the method successfully distilled six high-level AI use case categories and produced 107 validated evaluation scenarios, substantially enhancing the consistency, comparability, and real-world relevance of AI assessments.
This work addresses the limitations of existing AI governance frameworks, which rely on static metrics and post-hoc audits and thus lack the capacity for dynamic, real-time assessment of deployment readiness in high-risk systems—particularly regarding fairness discrepancies, threshold sensitivity, and remediation progress. To bridge this gap, the paper proposes the Operational AI Deployment Assurance (OADA) framework, which uniquely models governance uncertainty as an operational challenge within the deployment pipeline. OADA introduces mechanisms such as deployment assurance scores, readiness categorization, threshold stability zones, and governance escalation states to enable closed-loop, dynamic governance from evaluation to deployment. By integrating the Fairness Discrepancy Index (FDI) and FairRisk-FDI with threshold sensitivity analysis and repair-aware assurance evolution, OADA successfully identifies models deemed “compliant” by conventional metrics yet operationally unstable, offering a scalable deployment assurance paradigm for high-stakes domains like medical AI.