unit conversion

Harmonizing datasets by converting units, aligning coordinate systems, spatial/temporal resolutions and formats, and reconciling variable semantics and domain conventions across climate, land, ocean, cryosphere and socioeconomic sources onto a common analysis grid.

unitconversion

12-Month Skill Trend

Momentum and market value over time
Trending
Score
+20 in 12 mo
96
12 mo agoNow
Career
Value
+$12K in 12 mo
$42K/year
12 mo agoNow

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

To address the challenge of achieving FAIR interoperability and reusability for multi-source, heterogeneous scientific data—such as the RADx COVID-19 response data—in large-scale environments, this paper proposes a general-purpose, reproducible, and extensible data harmonization framework. Our approach introduces a novel harmonization paradigm based on parameterized primitive operations and automated execution tracing. It integrates a customizable data representation model, a configurable operation library, and mechanisms for transformation logging and dependency tracking—ensuring protocol reproducibility, process auditability, and transformation reusability. Evaluated in real-world deployment within the RADx Data Hub, the framework significantly lowers the barrier to entry for domain experts, improves harmonization efficiency and transparency, and enables high-quality cross-study analyses. This work provides a scalable, transferable technical pathway for FAIR-compliant data integration across diverse scientific domains.

Addresses challenges in aligning heterogeneous data setsFacilitates data harmonization for FAIR principles complianceSupports reproducible and scalable data integration processes

This work addresses a critical limitation in current Earth system foundation models—the absence of a unified training dataset that integrates heterogeneous, multi-source data spanning climate, land, ocean, cryosphere, infrastructure, hazards, and socioeconomic domains. To overcome this, the study constructs WorldTensor, a standardized dataset that harmonizes hundreds of environmental and socioeconomic variables into a common 0.25° spatial grid and annual temporal framework. Through techniques including regridding, rasterization of vector and point data, and temporal alignment, the resulting dataset adheres to Climate and Forecast (CF) metadata conventions and is distributed in NetCDF format. WorldTensor enables deep multimodal integration across natural and human systems, establishing a reproducible, high-quality benchmark for training and evaluating planetary-scale coupled models.

Earth system foundation modelsglobal training resourceharmonised dataset

Interactive Data Harmonization with LLM Agents

Feb 10, 2025
AS
Aécio Santos
🏛️ New York University | Federal University of Technology - Paraná

Standardizing multi-source heterogeneous clinical data remains challenging due to schema misalignment, terminological heterogeneity, and variability in data collection practices. Method: This paper proposes a novel interactive data harmonization paradigm powered by LLM-based agents, integrating domain expert knowledge with large language model reasoning. Through an interactive UI, users progressively construct harmonization pipelines, supported by core components including schema mapping, semantic alignment, and a standardized primitive library. Contribution/Results: Unlike end-to-end black-box approaches, our work introduces the first human-in-the-loop, incrementally generated, and on-demand reusable pipeline construction mechanism. Experiments demonstrate a ~70% reduction in manual coding effort, significantly improved harmonization consistency and reproducibility, and empirically validated effectiveness and generalizability on real-world clinical datasets.

Automating data harmonization processesCreating reusable data mapping pipelinesIntegrating datasets from diverse sources

10 Simple Rules for Improving Your Standardized Fields and Terms

Oct 21, 2025
RC
Rhiannon Cameron
🏛️ Simon Fraser University

Scientific data often suffers from poor discoverability, limited sharing, inefficient reuse, and high curation costs due to inadequate standardization of fields and terminology. To address these challenges, this paper proposes a FAIR-aligned standardization framework. Methodologically, it integrates structured vocabulary design, context-aware metadata modeling, and data homogenization strategies to systematically mitigate semantic noise and concept explosion. Crucially, it embeds ten actionable, principle-based rules into a dynamic, evolving data governance process. Empirical evaluation demonstrates that the framework significantly improves metadata quality and semantic consistency, reduces data management overhead, and enhances data findability, interoperability, and long-term reusability—thereby enabling robust, real-world implementation of the FAIR principles in scientific research settings.

Addressing challenges in standardizing research metadata vocabulariesOffering practical rules for FAIR-compliant metadata designProviding strategies to improve data findability and reusability

Spatial Data Science Languages: commonalities and needs

Mar 20, 2025
EP
E. Pebesma
🏛️ University of Münster | Charles University | Environmental Systems Research Institute, Inc. (Esri) | Adam Mickiewicz University | AIT Austrian Institute of Technology | Wherobots, Inc. | Deltares | Delft University of Technology | Norwegian Institute for Nature Research (NINA) | University of Leeds | Bochum University of Applied Sciences | University of Salzburg

This paper identifies and systematically analyzes common challenges in spatial data science across mainstream programming languages—R, Python, and Julia—including inconsistent spherical geometry modeling, ambiguous spatial/temporal semantics, conflation of intensive and extensive attributes, poor interoperability between data cube and vector formats, complex cross-package dependencies, and a persistent divide between GIS and physical modeling communities. Through multi-language ecosystem surveys, cross-community comparative analysis, and software engineering abstraction, we propose, for the first time, a cross-language semantic framework for spatial operations. The framework formally defines support types (point vs. block), specifies attribute-type constraints on operation validity, and refactors spherical Simple Features logic. We distill five foundational insights that establish a methodological basis and practical guidance for tool interoperability, pedagogical alignment, and open-source governance in spatial computing.

Addressing geometric and statistical challenges in spatial data handlingImproving cross-language tools and community diversity in spatial scienceStandardizing spatial data analysis across R, Python, and Julia

Latest Papers

What's happening recently
View more

This work addresses the challenges of data integration arising from heterogeneity in data schemas, value representations, and domain-specific conventions. To this end, it proposes a data harmonization framework that synergistically combines programmatic interfaces with natural language interaction. The framework integrates schema- and value-matching algorithms, AI-augmented reasoning, and composable harmonization primitives, enabling users to flexibly construct reusable harmonization pipelines via a Python API. Simultaneously, it offers domain experts a conversational interface to explore, validate, and refine results using natural language. By coupling automated matching with iterative user feedback, the system significantly enhances both the efficiency and usability of data harmonization, as demonstrated in two representative application scenarios.

data harmonizationheterogeneous dataintegrative analysis

Climate model wind field outputs suffer from coarse resolution and substantial systematic biases, limiting their utility for wind energy applications that demand spatial coherence, multivariate consistency, and robustness under future climate scenarios. This work proposes an unsupervised, multivariate downscaling and bias correction method based on the SerpentFlow framework, which operates without paired high- and low-resolution data. By disentangling large-scale circulation patterns from small-scale variability and integrating flow-matching generative modeling with an interpretable generative domain alignment strategy, the approach reconstructs wind fields that preserve spatial structure, maintain cross-variable consistency, and generalize to future climates. Experiments demonstrate that the method significantly outperforms conventional multivariate correction techniques in reproducing mean wind speed, extreme wind speed, and zonal–meridional components, while markedly enhancing spatial coherence, inter-variable fidelity, and stability across diverse climate scenarios.

bias correctionclimate modelsdownscaling

This study addresses the limitations of conventional “master-slave” paradigms in geospatial data fusion, which hinder symmetric utilization of multi-source information and impede cross-community, cross-scale collaboration, thereby constraining the full potential of geospatial data. To overcome this, the authors propose a “global–local loop” fusion framework that establishes a bidirectional feedback mechanism through symmetric interaction among remote sensing imagery (e.g., Sentinel), volunteered geographic information (e.g., OpenStreetMap), and deep learning models, thereby transcending traditional unidirectional dependency. Validation on representative tasks such as land cover mapping demonstrates that the proposed approach significantly enhances both generalizability and thematic performance, offering a novel pathway for synergistic multi-source data fusion.

community biasdata fusiongeospatial data

The rapid proliferation of Earth observation instruments has been hindered by the absence of a unified, reliable, and persistent source of metadata, impeding data discovery and interpretation. To address this challenge, this work proposes and implements the first open-source registry for Earth observation instruments—Awesome Earth Observation Instruments. The registry employs a lightweight core schema extensible through modular components to capture multidimensional instrument characteristics, including spectral, geometric, and data access properties. It integrates automated validation with human curation to ensure metadata quality. Hosted on GitHub, the registry leverages version control for continuous evolution and is designed for future API interoperability. This infrastructure substantially enhances the discoverability, interpretability, and analyzability of instrument metadata, thereby supporting more effective utilization of Earth observation data.

data discoveryEarth observationinstrument registry

This study addresses long-standing challenges in Brazilian surface meteorological observations, including heterogeneous data formats, inconsistent variable naming, and inadequate quality control, which have hindered reproducible research across multiple disciplines. We present a high-quality, hourly-resolution meteorological dataset spanning 2000–2025 from 616 stations, featuring an innovative pipeline that automatically parses and semantically aligns heterogeneous Portuguese-language source data. A novel diagnostic quality control framework is introduced, preserving original values while applying two-stage checks for physical plausibility and spatiotemporal consistency. The resulting dataset includes standardized variables, unified timestamps, comprehensive metadata, and supporting audit files—such as station inventories, daily precipitation summaries, and variable-level failure statistics—significantly enhancing data transparency and usability for climate, environmental, agricultural, and machine learning applications.

Brazildata harmonizationmeteorological data

Hot Scholars

DL

Daniel Levy

National Heart, Lung, and Blood Institute
GeneticsGenomicsEpidemiologyCardiovascular Disease