HorusEye: Language as Dynamic Attention for Emergency Visual Analysis

📅 2026-06-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the significant performance degradation of current vision-language models in emergency scenarios under visually degraded conditions such as smoke, haze, and thermal imaging, where both grounding and visual question answering capabilities deteriorate markedly and exhibit model-dependent responses to linguistic feedback. To systematically evaluate robustness, the authors introduce the RefCOCO-Degraded dataset and assess prominent multimodal large language models—including Gemini, Qwen2-VL, and BLIP-2—across four degradation types. The work reveals two key findings: the previously unreported “thermal imaging paradox” and BLIP-2’s heightened susceptibility to hallucination under image degradation. Experimental results demonstrate that iterative language feedback improves Gemini’s performance in thermal imaging by 47.3%, whereas Qwen2-VL suffers a 5.1% decline, and further uncover that standard cropping strategies can severely fail in thermal settings.
📝 Abstract
We introduce HorusEye, Language as Dynamic Attention for Emergency Visual Analysis. Our investigation followed five stages. The first one is benchmarking RefCOCO-Degraded, a dataset of 15,244 images (3,811 base images x 4 conditions: Clean, Fog, Smoke and Thermal) with systematic visual degradation. Through four research questions, we evaluate multiple VLMs (Gemini, Qwen2-VL, BLIP-2, LLaVA, Kosmos-2) across visual grounding the second stage, language feedback recovery the third one, health VQA tasks the fourth, and hallucination analysis the final stage. Our key finding is that language feedback effectiveness is model-dependent: Gemini achieves +47.3% improvement in thermal conditions through iterative language feedback, while Qwen2-VL shows -5.1% degradation under the same protocol. We also identify the'Thermal Paradox'where cropping strategies that improve RGB performance catastrophically fail in thermal imagery. Furthermore, BLIP-2 uniquely hallucinates more under degradation, making it unsuitable for emergency deployment
Problem

Research questions and friction points this paper is trying to address.

visual degradation
language feedback
visual grounding
hallucination
emergency visual analysis
Innovation

Methods, ideas, or system contributions that make the work stand out.

Dynamic Attention
Visual Degradation
Language Feedback
Thermal Paradox
Visual Grounding
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
A
Armel Yara
The Day Info