Aperture: Training-Free Multiscale Concept Bottlenecks for Remote Sensing

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limited interpretability of remote sensing models and the high training costs and poor unsupervised performance of existing concept bottleneck models by proposing a training-free multi-scale concept bottleneck framework. The method introduces a novel greedy quadtree routing mechanism in image space to localize fine-grained concepts, leverages pretrained multimodal large language models (MLLMs) for reliable concept scoring, and integrates global and local features to achieve fine-grained concept mapping. Additionally, a fine-grained concept dataset, SiFC, is constructed. Experimental results demonstrate that the proposed framework surpasses the strongest baseline on SiFC by over 10 percentage points in macro-F1 score, outperforms supervised approaches, and supports change detection without retraining.
📝 Abstract
While earth observation models have advanced substantially, they still lack interpretability. While concept-bottleneck models provide interpretability and expert interaction, they are either too expensive to train for the remote sensing domain or perform poorly without annotation. We posit that in expert domains like remote sensing, such training-free models require both fine details in both image and concept space. In image space, we propose a multiscale concept bottleneck using greedy quadtree routing to locate small concepts. In concept space, we replace contrastive vision language models with pre-trained MLLMs and present a way to get reliable concept scores from them. We introduce APERTURE that blends concept scores at the global image and native concept-scale level to give state-ofthe-art training-free model performance. To test these models, introduce SiFC, a fine-grained concept-centric dataset across three countries, with human-reviewed class-level concept maps. On SiFC, APERTURE outperforms the best training-free baselines by more than 10 percentage points in macro F1-score, and notably also outperforms supervised concept bottleneck models. Targeted component-removal tests examine whether concept scores respond to changes in visual evidence, while temporal experiments show that descriptor updates improve recognition of technological changes without retraining.
Problem

Research questions and friction points this paper is trying to address.

Remote Sensing
Interpretability
Concept Bottleneck Models
Training-Free
Fine-Grained Concepts
Innovation

Methods, ideas, or system contributions that make the work stand out.

training-free concept bottleneck
multiscale quadtree routing
multimodal large language models
remote sensing interpretability
fine-grained concept dataset
🔎 Similar Papers
2024-03-18International Journal of Computer VisionCitations: 48