Hazard or Anomaly? Evaluating VLMs for Understanding Dangers and Discrepancies

๐Ÿ“… 2026-07-17
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This work addresses a critical limitation in current vision-language models (VLMs) wherein safety assessments are often confounded by the conflation of โ€œanomalousโ€ and โ€œharmfulโ€ content, a shortcoming rooted in conventional binary safety labels. To overcome this, the study proposes the first fine-grained evaluation framework that explicitly decouples harm from anomaly. Through systematic testing of multiple state-of-the-art VLMs on two curated datasets and employing diverse prompting strategies, the authors reveal a pervasive tendency of models to infer danger primarily from scene atypicality rather than genuine harmfulness. This approach substantially enhances the discriminative power of safety reasoning evaluations, uncovering failure modes obscured by traditional binary judgments. The paper also introduces and publicly releases a dedicated dataset to foster more nuanced and robust safety assessment paradigms for multimodal systems.
๐Ÿ“ Abstract
Modern safety-critical systems increasingly rely on human-robot interaction to reduce disaster risk and support decision-making during emergencies. Vision-Language Models (VLMs) are promising for these settings because they can interpret complex scenes and communicate safety-relevant information, but they still require careful evaluation to ensure reliable safety reasoning. In particular, current evaluations often frame danger recognition as a binary decision (Safe/Unsafe), making it unclear whether a model is identifying true physical hazards or merely reacting to unusual scene elements. We address this limitation by introducing an explicit distinction between hazard and anomaly, and by separately recognizing hazardous and anomalous states. We evaluate several state-of-the-art VLMs across two datasets and multiple prompting strategies to test whether this distinction changes model behavior. Our results show that VLMs frequently misinterpret anomalousness as hazardousness, revealing an over-reliance on contextual irregularity as a proxy for danger. We further show that explicitly separating anomaly from hazard provides a more informative evaluation of VLM safety reasoning and exposes failure modes that binary safety judgments can obscure. Our public dataset is available on Roboflow https://app.roboflow.com/vlm-in-context-anomaly-and-hazard-detection/camera-ready-roman-ds.
Problem

Research questions and friction points this paper is trying to address.

hazard
anomaly
Vision-Language Models
safety reasoning
danger recognition
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vision-Language Models
hazard detection
anomaly detection
safety reasoning
evaluation framework
M
Murali Indukuri
Interactive Robotics and Language Lab, University of Maryland Baltimore County, Computer Science and Electrical Engineering Department, Baltimore, MD, USA.
M
Mohammad Eskandari
Interactive Robotics and Language Lab, University of Maryland Baltimore County, Computer Science and Electrical Engineering Department, Baltimore, MD, USA.
S
Sree Nitya Kollu
Interactive Robotics and Language Lab, University of Maryland Baltimore County, Computer Science and Electrical Engineering Department, Baltimore, MD, USA.
S
Stephanie Lukin
DEVCOM Army Research Laboratory, Adelphi, MD, USA.
Cynthia Matuszek
Cynthia Matuszek
Associate Professor, UMBC
roboticsnatural language groundingmachine learningknowledge representation