EgoSafe: A First-Person Mobile-Captured Benchmark for Visual Safety Understanding

📅 2026-07-29
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Existing visual safety evaluations predominantly rely on third-person videos and binary classification metrics, which are inadequate for assessing models’ causal reasoning capabilities in first-person, dynamic, and partially observable scenarios. To address this gap, this work introduces EgoSafe-Bench—the first benchmark specifically designed for first-person visual safety understanding—comprising 3,000 egocentric mobile-captured videos and 12,000 question-answer chains generated under a Hierarchical Reasoning Evaluation (HRE) protocol. This protocol compels models to construct coherent reasoning chains spanning feature anchoring, occlusion-aware inference, and intent prediction. Evaluations of leading large vision-language models—including Qwen3-VL, Gemini, and VideoLLaMA 3—reveal strong performance on descriptive tasks but significant deficiencies in causal reasoning and logical consistency, thereby demonstrating the benchmark’s effectiveness and challenge.
📝 Abstract
Reliable visual safety understanding in real-world scenarios demands more than just object recognition; it requires causal reasoning under epistemic uncertainty. While Large Vision-Language Models (LVLMs) demonstrate impressive semantic alignment on standard benchmarks, they often struggle to distinguish between superficial correlation and genuine forensic logic when grounded in the dynamic, partially observable nature of first-person experiences. Existing evaluations, dominated by third-person surveillance footage and binary classification metrics, fail to expose this cognitive gap. To address this, we introduce EgoSafe-Bench, a benchmark specifically designed to probe forensic reasoning in egocentric safety scenarios. It comprises 12,000 unique evaluation samples, generated by pairing each of the 3,000 video clips with a QA chain governed by our proposed Hierarchical Reasoning Evaluation (HRE) protocol. Unlike standard benchmarks, HRE mandates a rigorous reasoning trajectory from initial feature anchoring to blind-spot deduction and intent inference, thereby enforcing logical consistency and penalizing shortcut-based predictions.Extensive evaluations of state-of-the-art LVLMs (e.g., Qwen3-VL, Gemini, VideoLLaMA 3) reveal a significant perception-reasoning decoupling: models often achieve high descriptive scores but exhibit notable fragility in causal reasoning and logical closure. Our work provides both a challenging dataset and a systematic evaluation framework to foster the development of logically robust video understanding systems.
Problem

Research questions and friction points this paper is trying to address.

visual safety understanding
egocentric vision
causal reasoning
epistemic uncertainty
forensic logic
Innovation

Methods, ideas, or system contributions that make the work stand out.

egocentric vision
causal reasoning
visual safety understanding
hierarchical reasoning evaluation
first-person video benchmark