Score
Designs and implements methods to link hazard information (points, polygons, or rasters) with non‑spatial records or administrative units by geocoding, spatial joins, proximity/buffer matching, and distance calculations, producing exposure indicators and distance‑to‑hazard measures. Handles coordinate system transformations, scale and aggregation mismatches, and location uncertainty when matching across datasets to support comparative or aggregated spatial analysis.
To address the multifaceted requirements of scientific research and location-based services—particularly in geocoding accuracy, robustness, and semantic understanding—this paper systematically analyzes evolutionary drivers and deconstructs core functional modules, establishing for the first time an input–output requirements framework tailored to diverse application scenarios. We propose a novel multi-paradigm collaborative architecture integrating rule engines, information retrieval, named entity recognition, geographic knowledge graphs, and large language models (LLMs). Based on this, we formulate design principles and a technical roadmap for next-generation geocoding systems: extensibility, high robustness, and semantic awareness. Key contributions include identifying three LLM-driven breakthrough directions: context-aware address parsing, cross-modal spatial-semantic alignment, and dynamic knowledge-enhanced reasoning—providing a systematic methodology for both academia and industry.
This paper identifies and systematically analyzes common challenges in spatial data science across mainstream programming languages—R, Python, and Julia—including inconsistent spherical geometry modeling, ambiguous spatial/temporal semantics, conflation of intensive and extensive attributes, poor interoperability between data cube and vector formats, complex cross-package dependencies, and a persistent divide between GIS and physical modeling communities. Through multi-language ecosystem surveys, cross-community comparative analysis, and software engineering abstraction, we propose, for the first time, a cross-language semantic framework for spatial operations. The framework formally defines support types (point vs. block), specifies attribute-type constraints on operation validity, and refactors spherical Simple Features logic. We distill five foundational insights that establish a methodological basis and practical guidance for tool interoperability, pedagogical alignment, and open-source governance in spatial computing.
Unstructured, heterogeneous, and inconsistently spelled location descriptions in disaster databases (e.g., EM-DAT) impede subnational geocoding. Method: We propose the first fully automated, GPT-4o–driven geocoding workflow: large language models perform text cleaning and semantic parsing; cross-validated geographic matching integrates GADM, OpenStreetMap, and Wikidata to generate subnational coordinates with reliability scores. Contribution/Results: The method enables flexible, multi-hazard, cross-administrative mapping and introduces the first LLM-powered, multi-source trustworthy geolocation framework. Applied to EM-DAT records from 2000–2024, it successfully geocoded 14,215 disaster events and 17,948 unique locations at subnational resolution, achieving high precision. This significantly enhances spatial comparability, interoperability, and analytical utility of disaster data.
This study addresses the challenge of balancing disclosure risk and data utility in the release of geospatial statistical maps, where existing risk measures are often unstable due to the modifiable areal unit problem (MAUP). To overcome this limitation, the authors propose a novel disclosure risk metric that explicitly incorporates local population density and multi-unit spatial dependencies into its formulation. The resulting framework adaptively assesses disclosure risk across varying map resolutions and zoom levels, effectively mitigating MAUP-induced instability. Empirical validation on simulated datasets mimicking real-world business locations demonstrates that the proposed method consistently and accurately reflects disclosure risk under diverse spatial partitioning and scaling scenarios, thereby substantially enhancing the robustness and practical applicability of risk assessment in geospatial data dissemination.
To address scale misalignment and information loss in spatial data arising from aggregation or registration, this paper proposes a Bayesian decomposition framework that maps misaligned observations—including point patterns and aggregated counts—onto a continuous spatial domain, enabling uncertainty-aware inversion under four covariate scenarios. We introduce an INLA-driven iterative linearization integration algorithm and design three covariate field reconstruction strategies: Value Plugin, Joint Uncertainty, and Uncertainty Plugin—explicitly propagating uncertainty while maintaining robustness to model misspecification. The method integrates point processes, hierarchical modeling, and multi-source covariates (raster, polygon, and point). In landslide susceptibility mapping, it substantially improves spatial resolution and predictive reliability. Notably, even under covariate scarcity, the Uncertainty Plugin maintains high accuracy, outperforming conventional interpolation and deterministic inversion approaches.
This work addresses non-uniform, nonlinear spatial misalignment between post-disaster orthoimagery from small Unmanned Aerial Systems (sUAS) and prior building vector polygons—a critical yet previously unquantified challenge. Method: We conduct the first large-scale quantitative analysis across 51 orthoimages from nine disasters and 21,600 buildings, revealing an average translational error of 82 pixels and mean IoU of only 0.65; remarkably low angular and distance variances (0.4° and 0.45 px) confirm the absence of spatial consistency—invalidating the common assumption that linear transformations suffice, as often assumed in satellite remote sensing. We propose a rigorous validation framework integrating geometric metrics, IoU-based assessment, and GIS overlay comparison, and introduce the first publicly available benchmark dataset with manually refined ground-truth annotations. Contribution/Results: Our findings expose significant bias risks for downstream AI systems and establish a reproducible, paradigm-shifting foundation for sUAS georegistration research.
Existing spatial pattern matching methods are largely confined to two-dimensional space and struggle to handle three-dimensional entity matching involving elevation or height information in real-world scenarios. This work extends spatial pattern matching to 3D environments for the first time, introducing a general problem formulation and proposing a subgraph-matching-based algorithm that explicitly models distance relationships in three-dimensional space. To support empirical evaluation, the authors construct the first 3D spatial pattern matching dataset, integrating both synthetic data and real-world building structures from the city of Hamburg. Experimental results on this benchmark demonstrate the effectiveness of the proposed approach, establishing a foundational algorithmic framework and experimental platform for future research in 3D spatial pattern analysis.
This study addresses the challenge of identifying small-scale informal environmental health hazards—such as informal lead-acid battery recycling—that evade detection by satellite observations and official registries. To overcome reliance on conventional remote sensing and registration data, the authors propose a contextualized geospatial feature construction method integrating domain knowledge. Leveraging geographic information systems and machine learning, the approach is rigorously evaluated through five-fold cross-validation, matched controls, and an independent dataset across India and Bangladesh. The model significantly outperforms random urban controls in detecting 172 previously unseen informal recycling sites and maintains high specificity by distinguishing informal activity patterns even among 131 formal facilities. This work represents the first scalable, highly specific method for identifying informal environmental risk sources.
This study addresses the strong co-occurrence of flood and landslide hazards and their spatially heterogeneous relationships with environmental factors, which challenge conventional models in balancing regional specificity and generalizability. A dual-strategy modeling framework integrating spatial proximity and ecological zoning is proposed to map compound flood–landslide susceptibility and relative risk in Kerala, India, and Nepal. The approach partitions the study areas into 15 km contextual units and employs two strategies: proximity-gated cross-zone training (S1) and ecozone-gated within-zone constrained training (S2). Using random forest, CRITIC weighting, and SHAP interpretability, coupled with spatial hold-out validation and multi-metric evaluation, results show S1 outperforms in most metrics (e.g., AUC-ROC of 0.886 for floods in Nepal), while S2 preserves region-specific factor contributions. The resulting nine-class compound hazard risk maps achieve a consistency of 0.711 in Nepal, with exposure and vulnerability substantially reshaping high-risk priority zones.
Geospatial impact evaluations often grapple with ambiguity in defining the exposure units, timing, and intensity of interventions, particularly when multiple plausible exposure definitions exist. This study introduces the concept of “treatment geometry” as a foundational framework to systematically characterize the spatiotemporal footprint of interventions derived from Earth observation data. Centered on key trade-offs—including spatial resolution, temporal alignment, spillover effects, and boundary uncertainty—the framework provides diagnostic tools that enable researchers to identify which geometric definition choices are most critical for causal identification, rather than defaulting to a single methodological approach. Empirical applications to air pollution, wildfires, and forest policy demonstrate that this approach substantially enhances the credibility of causal inference and the rigor of empirical design.
This study addresses the mismatch between areal-aggregated geographic data and point-based spatial scan statistics, where representing regions by their centroids discards critical spatial information and reduces statistical power. To mitigate this limitation, the authors propose a simple yet scalable preprocessing strategy: uniformly sampling 20–50 points within each region’s geometry and distributing the region’s observed count equally among these points. This approach better preserves the underlying spatial distribution while remaining computationally tractable. Empirical evaluations demonstrate that the method substantially enhances the detection performance of spatial scan statistics on aggregated regional data across diverse scenarios. The authors advocate its adoption as a standard preprocessing step for analyzing areal-aggregated datasets in spatial anomaly detection tasks.