Absence is Presence: Understanding Visual Scene Negative Events Under Safety Cognitive Constraint

📅 2026-09-17
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究解决了安全关键领域中理解图像中缺失信息的问题,通过提出基于反事实重构和对比解码的框架来推断应存在但实际不存在的关键信息。
📝 Abstract
Traditional scene understanding focuses on affirmative information objectively present in images. However, in safety-critical domains, comprehending key information that should exist but is actually absent is vital for risk mitigation. To bridge this gap, we focus on visual scene negative captioning with safety as the cognitive constraint. The core challenge is to convert physical absence into semantic negative events. Existing vision-language models (VLMs) struggle with this process because affirmation bias suppresses negative reasoning, while limited mental filling capability and representation bias further hinder the inference of absent information. To address these challenges, we propose a negative captioning framework based on counterfactual reconstruction and contrastive decoding (CRCD). Inspired by human cognition, CRCD reformulates the task as counterfactual latent change captioning to bypass affirmation bias. It contrasts a synthesized safe expectation with reality to identify semantic omissions. To address limited mental filling, we design a dual-branch counterfactual reconstruction architecture. The amodal completion branch restores defective objects, while the functional association branch infers completely absent safety objects. Concurrently, a multi-condition representation learning mechanism is integrated to mitigate representation bias by projecting universal features onto predefined safety criteria subspaces, thereby capturing information across more dimensions. By decoding feature-level semantic residuals between the reconstructed scene prototype and raw input, CRCD bounds the non-existence search space and activates the decoder's negative logic. Extensive experiments validate the effectiveness of CRCD, establishing a high-performance baseline for this pioneering task.
Problem

Research questions and friction points this paper is trying to address.

visual scene negative events
safety cognitive constraint
affirmation bias
mental filling capability
representation bias
Innovation

Methods, ideas, or system contributions that make the work stand out.

counterfactual reconstruction and contrastive decoding (CRCD)
safety cognitive constraint
visual scene negative captioning
representation bias
multi-condition representation learning
Z
Zhiyun Jiang
Sichuan University, Chengdu, China
H
Hanyong Wang
Sichuan University, Chengdu, China
B
Binbin Liang
Sichuan University, Chengdu, China
Y
Yu Xie
Beijing Institute of Technology, Beijing, China
Menglong Yang
Menglong Yang
Sichuan University
Computer Vision
Wei Li
Wei Li
Sichuan University
Camera networks