HomeSafeBench: A Benchmark for Embodied Vision-Language Models in Free-Exploration Home Safety Inspection

📅 2025-09-28
📈 Citations: 0
Influential: 0
📄 PDF

career value

226K/year
🤖 AI Summary
Existing home safety inspection benchmarks suffer from two critical limitations: they rely on textual descriptions instead of authentic visual inputs, and employ only a single static viewpoint—severely undermining the evaluation of embodied perception and active exploration capabilities in vision-language models (VLMs). To address this, we introduce ESBench, the first benchmark specifically designed for evaluating embodied VLMs on home safety tasks. Built upon a high-fidelity simulation environment, ESBench generates first-person, dynamic image sequences enabling multi-view, free-exploration detection of five hazard categories: fire, electric shock, falling objects, tripping hazards, and child safety risks. By departing from conventional static or text-based paradigms, ESBench more faithfully captures the perception–decision closed loop in complex domestic environments. Experimental results reveal severe performance limitations of state-of-the-art VLMs on ESBench (best F1 score: 10.23%), highlighting fundamental deficiencies in fine-grained hazard recognition and exploration strategy learning.

Technology Category

Application Category

📝 Abstract
Embodied agents can identify and report safety hazards in the home environments. Accurately evaluating their capabilities in home safety inspection tasks is curcial, but existing benchmarks suffer from two key limitations. First, they oversimplify safety inspection tasks by using textual descriptions of the environment instead of direct visual information, which hinders the accurate evaluation of embodied agents based on Vision-Language Models (VLMs). Second, they use a single, static viewpoint for environmental observation, which restricts the agents' free exploration and cause the omission of certain safety hazards, especially those that are occluded from a fixed viewpoint. To alleviate these issues, we propose HomeSafeBench, a benchmark with 12,900 data points covering five common home safety hazards: fire, electric shock, falling object, trips, and child safety. HomeSafeBench provides dynamic first-person perspective images from simulated home environments, enabling the evaluation of VLM capabilities for home safety inspection. By allowing the embodied agents to freely explore the room, HomeSafeBench provides multiple dynamic perspectives in complex environments for a more thorough inspection. Our comprehensive evaluation of mainstream VLMs on HomeSafeBench reveals that even the best-performing model achieves an F1-score of only 10.23%, demonstrating significant limitations in current VLMs. The models particularly struggle with identifying safety hazards and selecting effective exploration strategies. We hope HomeSafeBench will provide valuable reference and support for future research related to home security inspections. Our dataset and code will be publicly available soon.
Problem

Research questions and friction points this paper is trying to address.

Evaluating embodied agents' home safety inspection capabilities using visual information
Addressing limitations of single static viewpoints in hazard detection tasks
Assessing VLM performance in identifying five common home safety hazards
Innovation

Methods, ideas, or system contributions that make the work stand out.

Dynamic first-person perspective images for evaluation
Free exploration in simulated home environments
Multiple dynamic perspectives for thorough inspection