Score
Designs and implements scalable preprocessing pipelines for large geospatial and satellite imagery, performing operations such as tiling and mosaicking, radiometric and geometric normalization, missing-tile detection and filtering, and preparation of batched inputs for models. Builds and optimizes data ingestion, storage and streaming components to handle high-throughput image I/O, parallelize transformations, and produce consistent, quality-controlled inputs for downstream analysis and machine learning.
This work addresses the unique challenges in Earth observation machine learning—such as georeferenced imagery, heterogeneous labels, and spatially aware sampling—and the lack of standardized tools that seamlessly integrate with mainstream deep learning frameworks. Building upon the TorchGeo library, this study presents the first systematic integration of geospatial data processing within the PyTorch ecosystem, offering an end-to-end, reproducible development paradigm. The methodology is demonstrated through a semantic segmentation case study on water bodies using Sentinel-2 imagery, showcasing a cohesive workflow that combines unified coordinate transformations, spatially aware samplers, pretrained models, and interactive Jupyter Notebook tutorials. The pipeline directly produces GeoTIFF prediction outputs suitable for downstream geospatial analysis. All code and tutorials are publicly released.
This study addresses the cumbersome deployment of remote sensing deep learning models within Geographic Information Systems (GIS) caused by format incompatibilities and computational disparities. We propose an open-source geospatial inference system featuring a novel decoupled architecture that separates multi-source heterogeneous models from arbitrary computing backends via unified interfaces to abstract underlying differences. Integrated as a QGIS plugin, the system enables automated image tiling, result reassembly, and vision-language model prompt injection, facilitating zero-code interactive real-time inference. Experimental results demonstrate that the system successfully executes parallel comparative evaluations of three heterogeneous models across cloud APIs, remote GPUs, and local CPUs, substantially improving efficiency for tasks such as agricultural parcel segmentation.
This study addresses land use/land cover (LULC) classification in remote sensing imagery by systematically benchmarking training performance of ResNet-50 across heterogeneous GPUs: Apple M3 Pro (integrated), NVIDIA RTX 3060 (consumer-grade), and Tesla T4 (cloud accelerator). We propose a lightweight, containerized training framework supporting cross-GPU deployment, integrated with Sentinel-2 and EuroSAT datasets, geospatial preprocessing, and an automated end-to-end training pipeline enabling reproducible execution. Our key contribution is the empirical validation—previously unreported—that freely available cloud GPUs and mainstream consumer hardware are viable for remote sensing deep learning. Relative to the M3 Pro, the RTX 3060 and T4 achieve up to 2× higher training throughput while maintaining >96% classification accuracy. The results provide a principled, cost-effective, scalable hardware selection guideline and engineering blueprint for resource-constrained geospatial AI applications.
This work addresses memory exhaustion and I/O bottlenecks in processing petabyte-scale image datasets—such as 1.4 PB electron microscopy volumes or 150 TB organ atlases—by introducing a streaming single-pass architecture based on a sweep execution model. The approach aligns disk reads with a one-dimensional sweep order and combines windowed operations with overlap-aware tiling to enable efficient processing under tight memory constraints. A domain-specific language (DSL) is designed to automatically optimize window sizes, fuse pipeline stages, and schedule multi-pass sweeps at compile time and runtime. The system supports Zarr, HDF5, and slice-based formats without requiring full-image residency in memory, achieving significantly higher throughput, near-linear I/O scaling, and predictable memory usage while seamlessly integrating with existing segmentation and morphological analysis toolchains.
为解决多尺度卫星图像合成问题,提出Genesis生成引擎,通过垂直超分辨率和水平外绘模型,实现跨尺度和空间的一致性。
Satellite imagery generates hundreds of terabytes of data daily, yet conventional compression methods apply uniform processing across entire images, struggling to balance downstream task requirements with compression efficiency. This work proposes a saliency-driven, region-adaptive compression approach that leverages saliency maps to guide spatially varying smoothing kernels during preprocessing, followed by standard lossy compression codecs such as JPEG. By integrating saliency-aware preprocessing with widely adopted compression standards, the method enables task-oriented variable bitrate allocation across image regions. It is the first to combine perceptual saliency with generic compression pipelines, significantly reducing storage and transmission overhead while preserving performance on downstream tasks—thereby overcoming the limitations inherent in globally uniform compression strategies.
This study addresses the lack of systematic best practices in large-scale Earth observation (EO) mapping, which often introduces errors during data preprocessing, model training, inference deployment, and validation, thereby compromising the reliability and scientific credibility of map products. To remedy this, we propose the first end-to-end best practice framework for EO mapping, encompassing the entire workflow from satellite data acquisition to operational map delivery. The framework integrates six core components: EO data infrastructure, preprocessing, machine learning dataset construction, uncertainty quantification, map production and dissemination, and independent validation. Emphasizing the interdependence of these stages, it embeds uncertainty quantification and independent validation as integral elements. By synergizing machine learning, distributed computing, and geospatial validation techniques, the framework establishes a reproducible and scalable mapping pipeline that substantially enhances the quality, consistency, and scientific rigor of EO-derived maps, supported by open-source resources to foster community adoption.
This work addresses the challenge that Python scripts authored by remote sensing scientists often lack scalability for large-scale satellite data processing. To bridge this gap, the authors propose an intelligent agent system that automatically translates existing Python geospatial workflows into efficient Apache Spark programs without requiring users to learn new frameworks. The system innovatively enhances the Scala-based RDPro library’s compatibility with large language models through structured API wrappers, function alias mapping, and an error-log-driven repair mechanism. Built upon LangGraph, it implements a staged pipeline for code generation and localized correction. Experiments on real-world geospatial workflows demonstrate that the approach correctly and efficiently processes massive remote sensing datasets, substantially improving scalability while preserving the original workflow semantics.
This work addresses the challenge of deploying lightweight remote sensing image segmentation models under stringent latency and energy constraints in on-orbit and edge computing scenarios. We propose a hardware-aware neural architecture search method that compresses the high-dimensional search space into two critical width variables. Leveraging CM-UNet as a teacher model, our approach integrates knowledge distillation with a low-fidelity regression surrogate model sampled on the Jetson Orin Nano to jointly predict segmentation accuracy and hardware costs, enabling rapid architecture selection. The method substantially improves search efficiency, achieving competitive mean Intersection-over-Union (mIoU) while significantly reducing latency and power consumption. Experimental results validate the effectiveness of the reduced-space regression strategy tailored for CNN-Mamba hybrid architectures in satellite-edge deployment settings.
本文通过利用NVIDIA Jetson平台的NVJPEG硬件加速单元和多实例方法,解决了边缘设备上JPEG格式数据解码速度慢的问题,提高了视觉模型推理效率。
Existing methods for satellite image generation struggle to effectively leverage diverse geospatial vector primitives—such as polygons, polylines, and points—to achieve high-fidelity, controllable synthesis. This work proposes the first unified spatial control framework based on a diffusion transformer architecture, incorporating a geometry-aware local attention mechanism that directly ingests arbitrary native vector primitives as conditioning inputs. By explicitly injecting geometric information, the model enables multi-granular spatial layout control. The approach consistently outperforms existing dense or sparse control methods across various vector conditions, achieving high-quality, controllable image generation with a single unified model. Furthermore, it significantly enhances performance on downstream tasks, including land cover segmentation, object detection, road extraction, and scene classification.