CogVis: Must Open-Vocabulary Change Detection Perceive the Scene Anew for Every Query?

📅 2026-08-06
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the instability and computational redundancy in open-vocabulary change detection caused by the tight coupling of temporal perception, semantic discrimination, and region verification. Inspired by human visual change perception mechanisms, we propose CogVis, a novel framework that introduces a cognitive memory mechanism to reformulate the task into a “perceive–remember–verify” paradigm. This decouples temporal perception from semantic decision-making, enabling cross-query sharing of change priors and dynamic calibration of class-specific thresholds. CogVis comprises three stages: Scene Change Perceptron, Semantic Memory Calibrator, and Adaptive Region Filter, integrating frozen bitemporal feature extraction with adaptive region filtering. The framework achieves state-of-the-art performance across seven benchmarks spanning semantic change detection, binary localization, and damage assessment, while improving inference throughput by 28.5%.
📝 Abstract
Earth-surface monitoring requires change detection models capable of recognizing arbitrary semantic categories. Open-Vocabulary Change Detection (OVCD) addresses this need. However, existing methods often entangle temporal perception, semantic discrimination, and region verification, causing unstable results and redundant computation. Inspired by human visual change perception, we propose CogVis, a cognitive memory-guided framework that reformulates OVCD as a perception-memory-verification paradigm. CogVis first employs a Scene Change Perceptron (SCP) to extract a reusable, category-agnostic change prior from frozen bi-temporal features, thereby decoupling temporal evidence from semantic category decisions. A Semantic Memory Calibrator (SMC) then compensates for category-dependent score shifts by dynamically estimating an image-query-specific decision threshold. Finally, an Adaptive Region Filter (ARF) filters connected candidates using learned semantic, temporal, and structural reliability. Experiments on seven benchmarks spanning semantic change detection, binary change localization, and building-damage assessment show that CogVis achieves state-of-the-art performance across all evaluated datasets. By sharing scene-level change perception, CogVis further avoids repeating category-agnostic temporal perception across queries and improves inference throughput by 28.50%.
Problem

Research questions and friction points this paper is trying to address.

Open-Vocabulary Change Detection
temporal perception
semantic discrimination
region verification
change detection
Innovation

Methods, ideas, or system contributions that make the work stand out.

Open-Vocabulary Change Detection
Cognitive Memory
Perception-Memory-Verification Paradigm
Change Prior Decoupling
Adaptive Region Filtering
🔎 Similar Papers
No similar papers found.