Score
Harmonizing datasets by converting units, aligning coordinate systems, spatial/temporal resolutions and formats, and reconciling variable semantics and domain conventions across climate, land, ocean, cryosphere and socioeconomic sources onto a common analysis grid.
This work addresses the fragmentation and lack of standardized interfaces in the ecosystem of geospatial foundation model embeddings, which severely hinder model comparison and reproducibility. We formalize, for the first time, the Earth embedding product ecosystem and propose a three-tier taxonomy grounded in data, tools, and value dimensions. Through a systematic analysis of interoperability barriers, we extend TorchGeo to develop a unified API that treats Earth embeddings as standardized geospatial datasets, enabling plug-and-play integration of heterogeneous, multi-source embedding products. This framework effectively decouples downstream analytical tasks from embedding engineering, substantially lowering the barrier to entry and promoting reproducibility, transparency, and fair benchmarking in remote sensing workflows.
To address the challenge of achieving FAIR interoperability and reusability for multi-source, heterogeneous scientific data—such as the RADx COVID-19 response data—in large-scale environments, this paper proposes a general-purpose, reproducible, and extensible data harmonization framework. Our approach introduces a novel harmonization paradigm based on parameterized primitive operations and automated execution tracing. It integrates a customizable data representation model, a configurable operation library, and mechanisms for transformation logging and dependency tracking—ensuring protocol reproducibility, process auditability, and transformation reusability. Evaluated in real-world deployment within the RADx Data Hub, the framework significantly lowers the barrier to entry for domain experts, improves harmonization efficiency and transparency, and enables high-quality cross-study analyses. This work provides a scalable, transferable technical pathway for FAIR-compliant data integration across diverse scientific domains.
This work addresses a critical limitation in current Earth system foundation models—the absence of a unified training dataset that integrates heterogeneous, multi-source data spanning climate, land, ocean, cryosphere, infrastructure, hazards, and socioeconomic domains. To overcome this, the study constructs WorldTensor, a standardized dataset that harmonizes hundreds of environmental and socioeconomic variables into a common 0.25° spatial grid and annual temporal framework. Through techniques including regridding, rasterization of vector and point data, and temporal alignment, the resulting dataset adheres to Climate and Forecast (CF) metadata conventions and is distributed in NetCDF format. WorldTensor enables deep multimodal integration across natural and human systems, establishing a reproducible, high-quality benchmark for training and evaluating planetary-scale coupled models.
Standardizing multi-source heterogeneous clinical data remains challenging due to schema misalignment, terminological heterogeneity, and variability in data collection practices. Method: This paper proposes a novel interactive data harmonization paradigm powered by LLM-based agents, integrating domain expert knowledge with large language model reasoning. Through an interactive UI, users progressively construct harmonization pipelines, supported by core components including schema mapping, semantic alignment, and a standardized primitive library. Contribution/Results: Unlike end-to-end black-box approaches, our work introduces the first human-in-the-loop, incrementally generated, and on-demand reusable pipeline construction mechanism. Experiments demonstrate a ~70% reduction in manual coding effort, significantly improved harmonization consistency and reproducibility, and empirically validated effectiveness and generalizability on real-world clinical datasets.
Scientific data often suffers from poor discoverability, limited sharing, inefficient reuse, and high curation costs due to inadequate standardization of fields and terminology. To address these challenges, this paper proposes a FAIR-aligned standardization framework. Methodologically, it integrates structured vocabulary design, context-aware metadata modeling, and data homogenization strategies to systematically mitigate semantic noise and concept explosion. Crucially, it embeds ten actionable, principle-based rules into a dynamic, evolving data governance process. Empirical evaluation demonstrates that the framework significantly improves metadata quality and semantic consistency, reduces data management overhead, and enhances data findability, interoperability, and long-term reusability—thereby enabling robust, real-world implementation of the FAIR principles in scientific research settings.
This paper identifies and systematically analyzes common challenges in spatial data science across mainstream programming languages—R, Python, and Julia—including inconsistent spherical geometry modeling, ambiguous spatial/temporal semantics, conflation of intensive and extensive attributes, poor interoperability between data cube and vector formats, complex cross-package dependencies, and a persistent divide between GIS and physical modeling communities. Through multi-language ecosystem surveys, cross-community comparative analysis, and software engineering abstraction, we propose, for the first time, a cross-language semantic framework for spatial operations. The framework formally defines support types (point vs. block), specifies attribute-type constraints on operation validity, and refactors spherical Simple Features logic. We distill five foundational insights that establish a methodological basis and practical guidance for tool interoperability, pedagogical alignment, and open-source governance in spatial computing.
This work addresses the challenges of data integration arising from heterogeneity in data schemas, value representations, and domain-specific conventions. To this end, it proposes a data harmonization framework that synergistically combines programmatic interfaces with natural language interaction. The framework integrates schema- and value-matching algorithms, AI-augmented reasoning, and composable harmonization primitives, enabling users to flexibly construct reusable harmonization pipelines via a Python API. Simultaneously, it offers domain experts a conversational interface to explore, validate, and refine results using natural language. By coupling automated matching with iterative user feedback, the system significantly enhances both the efficiency and usability of data harmonization, as demonstrated in two representative application scenarios.
Climate model wind field outputs suffer from coarse resolution and substantial systematic biases, limiting their utility for wind energy applications that demand spatial coherence, multivariate consistency, and robustness under future climate scenarios. This work proposes an unsupervised, multivariate downscaling and bias correction method based on the SerpentFlow framework, which operates without paired high- and low-resolution data. By disentangling large-scale circulation patterns from small-scale variability and integrating flow-matching generative modeling with an interpretable generative domain alignment strategy, the approach reconstructs wind fields that preserve spatial structure, maintain cross-variable consistency, and generalize to future climates. Experiments demonstrate that the method significantly outperforms conventional multivariate correction techniques in reproducing mean wind speed, extreme wind speed, and zonal–meridional components, while markedly enhancing spatial coherence, inter-variable fidelity, and stability across diverse climate scenarios.
This study addresses the limitations of conventional “master-slave” paradigms in geospatial data fusion, which hinder symmetric utilization of multi-source information and impede cross-community, cross-scale collaboration, thereby constraining the full potential of geospatial data. To overcome this, the authors propose a “global–local loop” fusion framework that establishes a bidirectional feedback mechanism through symmetric interaction among remote sensing imagery (e.g., Sentinel), volunteered geographic information (e.g., OpenStreetMap), and deep learning models, thereby transcending traditional unidirectional dependency. Validation on representative tasks such as land cover mapping demonstrates that the proposed approach significantly enhances both generalizability and thematic performance, offering a novel pathway for synergistic multi-source data fusion.
The rapid proliferation of Earth observation instruments has been hindered by the absence of a unified, reliable, and persistent source of metadata, impeding data discovery and interpretation. To address this challenge, this work proposes and implements the first open-source registry for Earth observation instruments—Awesome Earth Observation Instruments. The registry employs a lightweight core schema extensible through modular components to capture multidimensional instrument characteristics, including spectral, geometric, and data access properties. It integrates automated validation with human curation to ensure metadata quality. Hosted on GitHub, the registry leverages version control for continuous evolution and is designed for future API interoperability. This infrastructure substantially enhances the discoverability, interpretability, and analyzability of instrument metadata, thereby supporting more effective utilization of Earth observation data.
This study addresses long-standing challenges in Brazilian surface meteorological observations, including heterogeneous data formats, inconsistent variable naming, and inadequate quality control, which have hindered reproducible research across multiple disciplines. We present a high-quality, hourly-resolution meteorological dataset spanning 2000–2025 from 616 stations, featuring an innovative pipeline that automatically parses and semantically aligns heterogeneous Portuguese-language source data. A novel diagnostic quality control framework is introduced, preserving original values while applying two-stage checks for physical plausibility and spatiotemporal consistency. The resulting dataset includes standardized variables, unified timestamps, comprehensive metadata, and supporting audit files—such as station inventories, daily precipitation summaries, and variable-level failure statistics—significantly enhancing data transparency and usability for climate, environmental, agricultural, and machine learning applications.