From Pixels to Hierarchical Sequences: Quadtree Mask Encoding for Vision-Language Binary Change Detection

📅 2026-09-09
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
本文提出QUAKE-CD框架,通过将变化检测结果编码为语法约束的四叉树序列,改进了视觉-语言模型在像素级变化检测中的表现,特别是在处理小而分散的变化时。
📝 Abstract
Dense change detection in remote sensing requires vision-language models (VLMs) to compare bi-temporal images and generate accurate pixel-level masks. Existing VLMs are largely confined to change captioning outputs, and the few that produce pixel-level masks still rely on external decoders or flat text-as-mask serialization, which are less effective for small and fragmented changes. We introduce QUAKE-CD, a framework that recasts dense change prediction as syntax-verifiable structured generation. QUAKE-CD represents binary change masks as grammar-constrained quadtree token sequences, making the masks compact, syntactically checkable, and deterministically decodable within an autoregressive generation space. We further construct QUAKE-CoT, which pairs these sequences with chain-of-thought traces grounded in visual evidence, and jointly optimizes textual reasoning and spatial dense prediction through a progressive curriculum followed by grammar-gated dual-reward RL. On QUAKE-CoT, QUAKE-CD achieves 78.31% accumulated F1, outperforming decoder-based and flat text-as-mask VLMs while producing more faithful bi-temporal reasoning.
Problem

Research questions and friction points this paper is trying to address.

dense change detection
vision-language models
pixel-level masks
external decoders
text-as-mask serialization
Innovation

Methods, ideas, or system contributions that make the work stand out.

QUAKE-CD
quadtree encoding
syntax-verifiable structured generation
grammar-gated dual-reward RL
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
X
Xiao An
Wuhan University
R
Ruikang Zhang
Peking University
C
Chen Zhong
Wuhan University
Xuli Shen
Xuli Shen
Fudan University
Computer Vision
J
Jiaxing Sun
Shanghai Artificial Intelligence Laboratory
Jiang Wu
Jiang Wu
Shanghai Artificial Intelligence Laboratory
large language modelvision language model
Wei He
Wei He
Professor, Wuhan University
remote sensingcomputer vision