🤖 AI Summary
This study addresses the significant performance degradation of current vision-language models in emergency scenarios under visually degraded conditions such as smoke, haze, and thermal imaging, where both grounding and visual question answering capabilities deteriorate markedly and exhibit model-dependent responses to linguistic feedback. To systematically evaluate robustness, the authors introduce the RefCOCO-Degraded dataset and assess prominent multimodal large language models—including Gemini, Qwen2-VL, and BLIP-2—across four degradation types. The work reveals two key findings: the previously unreported “thermal imaging paradox” and BLIP-2’s heightened susceptibility to hallucination under image degradation. Experimental results demonstrate that iterative language feedback improves Gemini’s performance in thermal imaging by 47.3%, whereas Qwen2-VL suffers a 5.1% decline, and further uncover that standard cropping strategies can severely fail in thermal settings.
📝 Abstract
We introduce HorusEye, Language as Dynamic Attention for Emergency Visual Analysis. Our investigation followed five stages. The first one is benchmarking RefCOCO-Degraded, a dataset of 15,244 images (3,811 base images x 4 conditions: Clean, Fog, Smoke and Thermal) with systematic visual degradation. Through four research questions, we evaluate multiple VLMs (Gemini, Qwen2-VL, BLIP-2, LLaVA, Kosmos-2) across visual grounding the second stage, language feedback recovery the third one, health VQA tasks the fourth, and hallucination analysis the final stage. Our key finding is that language feedback effectiveness is model-dependent: Gemini achieves +47.3% improvement in thermal conditions through iterative language feedback, while Qwen2-VL shows -5.1% degradation under the same protocol. We also identify the'Thermal Paradox'where cropping strategies that improve RGB performance catastrophically fail in thermal imagery. Furthermore, BLIP-2 uniquely hallucinates more under degradation, making it unsuitable for emergency deployment