A Framework for Evaluating Vision-Language Model Safety: Building Trust in AI for Public Sector Applications

πŸ“… 2025-02-22
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
Public-sector deployment of vision-language models (VLMs) faces a critical security trust bottleneck due to the lack of standardized, quantitative safety evaluation frameworks tailored to governmental applications. Method: We propose the first VLM security quantification framework specifically designed for public administration scenarios. It introduces a novel Vulnerability Scoreβ€”a unified metric capturing performance degradation under both stochastic perturbations (Gaussian, salt-and-pepper, uniform noise) and adversarial attacks (e.g., FGSM). Further, we develop a composite noise-patch injection method coupled with saliency pattern analysis to precisely localize model vulnerabilities. Contribution/Results: Experiments demonstrate strong correlation (ρ > 0.92) between our Vulnerability Score and human expert security assessments. Robustness threshold modeling enables reproducible, quantifiable safety admission criteria. This work establishes both theoretical foundations and practical standards for AI governance in the public domain.

Technology Category

Computer Vision: Adversarial Attacks & RobustnessNatural Language Processing: Safety and RobustnessMachine Learning: Adversarial Learning & Robustness

Application Category

Security and Privacy: Large-scale security measurementsSearch and Retrieval-Augmented AI: Web evaluation methodologies and metricsUser Modeling, Personalization and Recommendation: Attacks and countermeasures in recommendation systems
πŸ“ Abstract
Vision-Language Models (VLMs) are increasingly deployed in public sector missions, necessitating robust evaluation of their safety and vulnerability to adversarial attacks. This paper introduces a novel framework to quantify adversarial risks in VLMs. We analyze model performance under Gaussian, salt-and-pepper, and uniform noise, identifying misclassification thresholds and deriving composite noise patches and saliency patterns that highlight vulnerable regions. These patterns are compared against the Fast Gradient Sign Method (FGSM) to assess their adversarial effectiveness. We propose a new Vulnerability Score that combines the impact of random noise and adversarial attacks, providing a comprehensive metric for evaluating model robustness.
Problem

Research questions and friction points this paper is trying to address.

Evaluate VLM safety vulnerabilities
Quantify adversarial risks in VLMs
Propose comprehensive Vulnerability Score
Innovation

Methods, ideas, or system contributions that make the work stand out.

Quantifies adversarial risks in VLMs
Analyzes noise impact on model performance
Introduces Vulnerability Score for robustness