🤖 AI Summary
This study addresses the cumbersome deployment of remote sensing deep learning models within Geographic Information Systems (GIS) caused by format incompatibilities and computational disparities. We propose an open-source geospatial inference system featuring a novel decoupled architecture that separates multi-source heterogeneous models from arbitrary computing backends via unified interfaces to abstract underlying differences. Integrated as a QGIS plugin, the system enables automated image tiling, result reassembly, and vision-language model prompt injection, facilitating zero-code interactive real-time inference. Experimental results demonstrate that the system successfully executes parallel comparative evaluations of three heterogeneous models across cloud APIs, remote GPUs, and local CPUs, substantially improving efficiency for tasks such as agricultural parcel segmentation.
📝 Abstract
Applying deep learning models to satellite imagery from within geographic information systems (GIS) remains high-friction for remote sensing practitioners. Models arrive in incompatible formats and target different compute environments, from local workstations to serverless cloud services. As a result, every evaluation demands custom deployment, tiling, and georeferencing code before a single prediction reaches the analyst's map. This friction discourages systematic comparison in a domain where model choice directly affects operational outcomes such as field delineation, crop monitoring, and disaster response. We present Anaximander, an open-source system that unifies model source and compute location choice behind one interactive interface. The system's backend is an inference server that loads models from multiple commonly-used sources and serves them on any accessible compute backend. The server provides session management and model caching, and streams results back per tile. The backend is paired with a QGIS plugin that drives tiling, result reassembly, georeferencing, and real-time per-tile status visualization. An additional user-interface path injects layer legends as prompts into vision-language models. We demonstrate the system in a code-free side-by-side comparison of three heterogeneous models on an agricultural field delineation task: gpt-image-1 via a cloud API, Segment Anything Model 3 (SAM3) on a remote GPU, and DelineateAnything on a local CPU. The inference backend and protocol are open-source and available at https://github.com/microsoft/nxmndr.