Benchmarking MLLMs via Cognitive Expected Scene Graph for Safety-Critical Visual Negation Understanding

📅 2026-09-17
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过构建安全认知下的场景否定理解任务和提出认知预期场景图评分方法,解决多模态大语言模型在处理视觉否定理解时的不足。
📝 Abstract
True machine intelligence requires transcending passive pixel registration to master top-down functional reasoning over absent information via visual negation understanding. However, unconstrained visual negation paradigms remain overly open-ended, and pervasive affirmation bias causes both existing Multi-Modal Large Language Models (MLLMs) and evaluation metrics to fail under negative semantics. To solve these intertwined challenges systematically, we first anchor the boundaries of negation reasoning within specific cognitive goals. Specifically, by focusing on safety as a highly pragmatic and critical cognitive dimension, we define the task of \textbf{S}cene \textbf{N}egation \textbf{U}nderstanding under \textbf{S}afety Cognition (\textbf{SNUS}). Under this framework, we construct a high-fidelity negative caption dataset mapping dense assertions of localized hazards. Concurrently, we propose the Cognitive Expected Scene Graph (CESG) Score, a structure-grounded, polarity-aware evaluation metric. Extensive experiments demonstrate that while current models struggle on the task, traditional metrics completely collapse under semantic reversals. Conversely, our framework delivers a solid benchmark for SNUS, providing a rigorous foundation to advance risk-aware situational comprehension and counterfactual cognition.
Problem

Research questions and friction points this paper is trying to address.

Visual Negation Understanding
Safety-Critical
Affirmation Bias
Multi-Modal Large Language Models
Evaluation Metrics
Innovation

Methods, ideas, or system contributions that make the work stand out.

Scene Negation Understanding under Safety Cognition (SNUS)
Cognitive Expected Scene Graph (CESG) Score
visual negation understanding
🔎 Similar Papers
Z
Zhiyun Jiang
Sichuan University, Chengdu, China
H
Hanyong Wang
Sichuan University, Chengdu, China
B
Binbin Liang
Sichuan University, Chengdu, China
Y
Yu Xie
Beijing Institute of Technology, Beijing, China
Menglong Yang
Menglong Yang
Sichuan University
Computer Vision
Wei Li
Wei Li
Sichuan University
Camera networks