🤖 AI Summary
This study addresses the poor reproducibility, lack of auditability, and reliance on manual narratives in AI safety assessments by proposing a deterministic, auditable framework. The framework standardizes heterogeneous engineering evidence into control identifiers mapped to technical-level risks, generates executable assessment functions by compiling MITRE ATLAS rules, supports repeated evaluations via versioned policy objects, and incorporates formal verification to ensure logical consistency and semantic correctness. Experiments across five open-source projects demonstrate that the framework effectively quantifies risk variations before and after hardening interventions. Results confirm that strengthened controls reduce attack feasibility while precisely revealing residual risks arising from missing core safeguards.
📝 Abstract
Artificial intelligence systems are increasingly deployed in high impact and safety critical settings, yet security assessment remains difficult to reproduce and defend under audit. Existing approaches often rely on narrative checklists or assessor driven scoring, and they lack an explicit, machine evaluable mapping from observable engineering artefacts to stable technique level outcomes. We present an evidence driven AI security assessment framework that operationalises assessment as a deterministic decision function. The framework normalises heterogeneous artefacts into a project independent Control ID taxonomy scored on a bounded four level ordinal scale, compiles technique level predicates from a pinned MITRE ATLAS snapshot via an explicit mitigation to control mapping, and outputs technique indexed feasibility and impact levels with traceable links back to the triggering evidence. We package all normative choices as a versioned assessment policy object to support repeatable reassessment across snapshots. To ensure semantic correctness, we formally verify boundedness, totality, ordered semantic consistency, and monotonicity of the compiled evaluator over the full declared score domain. We evaluate the framework on five public open source AI projects pinned to explicit repository snapshots, quantify before and after changes under a unified hardening intervention, and validate responsiveness to real engineering changes through fork based implementations of Software Bill of Materials (SBOM) generation and Continuous integration (CI) security scanning gates. Results show consistent downward shifts in feasibility profiles under strengthened observable controls, while worst case residual feasibility persists when technique specific core controls remain absent from the evidence scope.