Score
Designing and operationalizing strategies, governance, and engineering controls to reduce ethical, safety, and dual-use risks—e.g., participant protections, regulated deployment sandboxes, and operational procedures for degraded or reconnecting systems.
This work addresses the current lack of a systematic understanding of the capabilities of AI sandboxes in ensuring safety, security, and regulatory compliance, particularly within physical AI and cyber-physical systems. It proposes the first unified, assurance-oriented framework for AI sandboxes, introducing a formal boundary definition, a comprehensive sandbox taxonomy, a threat model targeting the assurance mechanisms themselves, and a quantifiable evaluation methodology spanning six dimensions—including fidelity and controllability. Through formal modeling, threat analysis, and multi-case validation, the study clarifies what aspects of AI behavior can be effectively tested in sandboxes, which risk categories can be meaningfully controlled, and what forms of evidence such environments can generate to support safety and compliance claims, thereby establishing foundational tools for trustworthy AI verification.
Conventional open-source risk management overrelies on technical tools, failing to address systemic risks—including upstream “silent fixes,” community conflicts, and sudden license changes—resulting in governance blind spots. Method: This paper proposes a strategic open-source risk governance framework centered on the interaction between external threats and internal vulnerabilities, shifting from tactical response to proactive, strategic prevention. It innovatively introduces a Strategic Objective Matrix and a dual-risk taxonomy, yielding an “Object–Threat–Vulnerability–Mitigation” decision model; integrates grounded theory, strategic mapping, and capability-building principles to support organization-level governance decisions. Contribution/Results: Validated by three domain experts and applied in real-world case studies, the framework significantly enhances risk analytical capability and enables enterprises to establish a systematic, immunizing mechanism against open-source risks.
Prior to large-scale AGI deployment, misuse and goal misalignment represent two critical safety risks. Method: We propose a dual-track collaborative defense framework: (1) at the model level, integrating amplified supervision, robust training, interpretability analysis, and uncertainty modeling; and (2) at the system level, implementing multi-tiered access control and real-time behavioral monitoring. Contribution/Results: This work is the first to systematically categorize and prioritize four risk types—misuse, misalignment, mistakes, and structural flaws—focusing explicitly on the former two. It introduces the first verifiable safety case framework for AGI, treating interpretability and uncertainty estimation as proactive enablers of safety assurance. The resulting end-to-end methodology comprehensively covers capability identification, safety hardening, dynamic monitoring, and failure containment—thereby enabling high-assurance, auditable, and formally verifiable AGI safety engineering.
Industrial operational technology (OT) systems—prioritizing functionality over security—are frequently misconfigured and exposed to the public Internet, posing severe cyber-physical risks. Method: We propose a comprehensive framework integrating cyberspace mapping, protocol fingerprinting, firmware version analysis, and a novel automated HMI/SCADA interface screenshot recognition technique to systematically assess global OT exposure. Our methodology correlates findings with vulnerability databases (e.g., NVD, ICS-CERT) and geolocation data across protocols, vendors, software, and regions. Contribution/Results: We identify nearly 70,000 publicly exposed OT devices, predominantly in North America and Europe; many run outdated firmware containing known critical vulnerabilities and remain unpatched for extended periods. Crucially, our interface-based analysis uncovers multiple previously undocumented unauthorized access paths—enabling the first large-scale, visually grounded quantification of real-world industrial attack surfaces and delivering actionable, operationally relevant insights for risk mitigation.
This paper addresses the conceptual conflation of “oversight” and “control” in AI safety governance, systematically distinguishing their distinct objectives, operational mechanisms, and temporal scopes. Through a critical cross-disciplinary literature review—and integrating insights from Responsible AI maturity models and risk governance theory—it develops a theoretically rigorous yet policy-actionable analytical framework, introducing the first AI Oversight Maturity Model (AI-OMM). The model identifies critical boundary conditions for oversight failure and establishes a structured, conditional system for assessing the feasibility of meaningful human oversight. Key contributions include: (1) clarifying the normative distinction between oversight and control; (2) diagnosing design gaps and contextual limitations in current oversight mechanisms; and (3) providing regulators, auditors, and developers with a practical tool to evaluate oversight effectiveness, detect capability gaps, and guide technical alignment with governance requirements.
Energy-sector industrial control systems (ICS) exhibit insufficient security resilience and overreliance on reactive, post-incident remediation. Method: This paper proposes a layered, implementable Security-by-Design (SbD) framework and a deployable set of security requirements tailored to critical infrastructure. Integrating systems engineering, ICS-specific security architecture, organizational behavior principles, and continuous monitoring, the approach spans the entire lifecycle—design, development, deployment, and operations—while ensuring alignment with IEC 62443 and NIST SP 800-82. Contribution/Results: It represents the first systematic, end-to-end operationalization of SbD in energy ICS contexts, enabling a paradigm shift from passive incident response to inherent, “native immunity.” The resulting scalable, auditable, and standards-coordinated SbD implementation guide supports the development of high-assurance, resilient, and sustainably evolvable cybersecurity ecosystems.
This study addresses the inadequacy of current IT compliance–oriented cybersecurity policies in safeguarding the physical safety of cyber-physical systems, as digital failures often precipitate real-world harm. By coding 292 critical infrastructure policies (2000–2025) and aligning them with the NIST SP 800-160 Vol. 2 resilience lifecycle, the research reveals a significant misalignment between prevailing policy approaches—overreliant on IT control catalogs during resistance and recovery phases—and actual physical risks. The work proposes a modernized “duty of reasonable care” standard centered on hazard-specific traceability, structured assurance cases, and cyber resilience engineering. It identifies three critical disconnects: misaligned delegation of standards, reduction of recovery mechanisms to mere incident reporting, and uneven sectoral adaptability. The study further outlines a viable pathway for federal policy that integrates engineering implementation with targeted incentives.
This study addresses the inefficiencies in SaaS onboarding within regulated enterprises, where siloed security and compliance controls—spanning third-party risk management, cybersecurity, identity and access management, and disaster recovery—often result in process delays, redundant assessments, and ambiguous accountability. To overcome these challenges, this work proposes an end-to-end, control-driven SaaS onboarding framework that integrates multi-domain controls into a unified lifecycle model encompassing requirement intake, architectural validation, identity design, resilience assessment, and post-deployment governance. By leveraging cross-domain control mapping, phased process modeling, and governance checklists, the framework codifies key design patterns such as secure connectivity, federated identity, least-privilege access, and shared-responsibility disaster recovery. Empirical implementation demonstrates that the approach significantly reduces onboarding friction, enhances audit traceability, and strengthens both the security posture and operational resilience of SaaS platforms.
This study addresses the systemic risks posed by on-premises AI coding agents, whose autonomous modifications to code and infrastructure may lead to severe organizational or societal harm due to challenges in timely constraint, auditing, or reversal. For the first time, it systematically applies three systems safety methodologies—STECA, STPA, and FRAM—to model risks in cutting-edge laboratory settings from multiple perspectives. The analysis reveals critical blind spots in current AI governance frameworks, particularly concerning unverifiable accountability, control failure caused by intervention delays, and weakened safeguards due to operational drift. The work underscores the necessity of integrating model-level evaluations with system-level hazard analysis, offering a crucial complementary pathway for robust AI risk management.
Existing decision graph approaches struggle to effectively model safety and security requirements in adaptive systems. This work proposes an extended decision graph modeling language that, for the first time, incorporates a safety-event dimension within a sustainability-driven framework and enables unified, synergistic modeling of safety, security, and sustainability through multi-granular “safety modes.” The approach supports formal specification and divide-and-conquer fine-grained management of relevant scenarios across the entire system lifecycle. Experimental evaluation on an industrial collaboration use case demonstrates that the proposed extension more accurately captures complex safety and security scenarios, significantly enhancing the overall modeling capability for adaptive systems.
Current AI incident governance frameworks lack consistency in defining, categorizing, monitoring, and reporting incidents, which constrains the depth and accuracy of post-deployment failure analysis. This study addresses this gap through a systematic literature review and comparative analysis across multiple governance frameworks, thereby identifying and synthesizing key inconsistencies that span existing mechanisms. The work reveals systemic deficiencies in data collection practices, classification logics, and analytical rigor, and elucidates critical misalignments among core governance components. By clarifying these structural disconnects, the research establishes a theoretical foundation and proposes a coordinated pathway toward a unified, standardized framework for AI incident governance.