process climate data

Designs and implements workflows and algorithms to ingest, clean, harmonize, and transform climate observations and model outputs into analysis‑ready datasets. This includes merging heterogeneous climate records, correcting biases and inhomogeneities, computing daily and regional statistics, and producing consistent time series for trend and other analyses.

processclimatedata

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.03
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Machine Learning Workflows in Climate Modeling: Design Patterns and Insights from Case Studies

Sep 30, 2025
TZ
Tian Zheng
🏛️ Columbia University | NSF STC Learning the Earth with AI and Physics (LEAP) | Data Science Institute | Jet Propulsion Laboratory | California Institute of Technology

This paper addresses core challenges in applying machine learning (ML) to climate modeling—including physical inconsistency, difficulties in multi-scale coupling, data sparsity, poor generalization robustness, and low integration with scientific workflows—by proposing a synergistic ML workflow framework that jointly leverages physical priors and simulation data. Methodologically, it integrates surrogate modeling, ML-driven parameterization, probabilistic programming, simulation-based inference, and physics-informed transfer learning, prioritizing interpretability and scientific rigor. Key contributions include: (i) distilling cross-task workflow design patterns; (ii) establishing a scientific ML practice framework that supports transparent development, critical evaluation, and reliable integration; and (iii) significantly enhancing model trustworthiness and interdisciplinary reproducibility while lowering technical barriers to deep integration between data science and climate modeling.

Developing ML workflows for climate modeling applicationsEnsuring physical consistency in machine learning climate modelsIntegrating scientific workflows with data-driven climate approaches

ClimateSOM: A Visual Analysis Workflow for Climate Ensemble Datasets

Aug 08, 2025
YK
Yuya Kawakami
🏛️ University of California, Davis | Scripps Institution of Oceanography | University of California, San Diego

Interpreting spatiotemporal variability patterns across climate model ensembles remains challenging due to high dimensionality and structural complexity. Method: This paper proposes a visualization analytics workflow integrating Self-Organizing Maps (SOM) with Large Language Models (LLMs). SOM reduces dimensionality and clusters high-dimensional climate time series while preserving spatial structure; LLMs semantically interpret clustering outcomes, generating scientifically grounded, human-readable descriptions of climate patterns. The integrated system supports interactive exploration of variability magnitude, spatial configurations, and inter-model grouping relationships. Contribution/Results: Evaluated on precipitation projections over California and the U.S. Pacific Northwest, the method enables accurate identification of dominant variability modes, reveals model consensus and divergence, and produces expert-validated, interpretable insights. This work represents the first deep integration of LLMs into an SOM-driven climate ensemble analysis pipeline, significantly enhancing cognitive efficiency and interpretability of complex ensemble variability structures.

Analyze variability in climate ensemble model projectionsIntegrate LLMs for interpreting climate data variabilityVisualize patterns and clusters in ensemble model runs

Taking the Garbage Out of Data-Driven Prediction Across Climate Timescales

Aug 09, 2025
JC
Jason C. Furtado
🏛️ University of Oklahoma | University of Maryland, College Park | Colorado State University | University of Lausanne | Ocean Associates, Inc. | NOAA/NWS/NCEP/Climate Prediction Center | Northern Illinois University | University of California, Davis | Instituto Geofísico del Perú | NOAA/GFDL | Macquarie

AI/ML models for climate prediction suffer from degraded skill and low credibility due to poor input data quality—including outliers, nonstationarity, and inadequate handling of spatiotemporal dependencies. Method: This study systematically identifies critical data preprocessing factors and proposes the first standardized AI/ML preprocessing protocol tailored for cross-timescale climate forecasting (subseasonal to decadal). It innovatively integrates standardized anomaly construction, nonstationary time-series correction, robust extreme-value handling, and modeling of complex-distribution variables. Multi-case empirical evaluations quantify the differential impacts of preprocessing strategies on prediction error, uncertainty quantification, and interpretability. Contribution/Results: The generalizable protocol significantly enhances model robustness and physical consistency, reducing bias by 15–30%. It establishes a foundational framework for standardized, transparent, and trustworthy climate AI applications.

Addressing data preprocessing impact on climate prediction modelsEstablishing protocols for AI/ML climate data preprocessingImproving robustness and transparency in climate prediction studies

This work proposes a machine learning–oriented paradigm for weather forecasting that reimagines the traditionally complex and closed operational systems to meet the demands of efficiency, openness, and collaboration in the era of artificial intelligence. By integrating agent-driven software engineering, open compressed data formats, shared validation workflows, interactive computing environments, and generative AI techniques, the framework systematically transforms model development, data utilization, computational management, and service delivery. Designed to equip meteorological and climate centers with future-ready infrastructure, it establishes robust data governance mechanisms, quality assurance protocols, and pathways for workforce skill transformation. The approach maintains scientific rigor while substantially enhancing the accessibility, efficiency, and interactivity of forecasting services.

digital transformationforecasting value chainmachine learning

CLIMATEAGENT: Multi-Agent Orchestration for Complex Climate Data Science Workflows

Nov 25, 2025
HK
Hyeonjae Kim
🏛️ The Hong Kong University of Science and Technology

Climate science urgently requires automated analytical frameworks capable of handling large-scale, heterogeneous data, yet existing general-purpose LLM agents and static scripts lack domain specificity and dynamic collaboration capabilities. To address this, we propose Climate-Agent: the first dynamic multi-agent framework tailored for climate data science. It enables end-to-end automation—from problem understanding and data acquisition to code generation and report synthesis—via API-aware task decomposition, a self-correcting execution loop, and coordinated operation of four specialized agent types (Orchestration, Planning, Data, and Coding). We introduce Climate-Agent-Bench-85, the first real-world benchmark comprising 85 complex climate science tasks. On this benchmark, Climate-Agent achieves a 100% task completion rate and an average report quality score of 8.32, significantly outperforming baseline methods including GitHub Copilot and GPT-5.

Automating climate data workflows across heterogeneous datasets and complex questionsEnabling reliable end-to-end climate science analysis with dynamic API awarenessOvercoming limitations of generic LLM agents and static scripting pipelines

Latest Papers

What's happening recently
View more

This study addresses the absence of unified evaluation criteria for climate models projecting regional temperature and humidity in the 2050s by proposing a three-tier probabilistic benchmark framework encompassing physical plausibility, observational fidelity, and paleoclimate extrapolation. Methodologically, it leverages post-2015 withheld data to achieve genuine out-of-sample validation and employs perfect-model experiments to quantify information value. The assessment integrates multiple dimensions through the Continuous Ranked Probability Score (CRPS), energy conservation diagnostics, and distributional consistency checks. By open-sourcing all code and datasets, this work establishes an equitable comparison platform for multi-paradigm climate models and provides a quantifiable evaluation system for tracking progress in climate projections.

benchmarkingclimate model evaluationout-of-distribution generalization

This study addresses long-standing challenges in Brazilian surface meteorological observations, including heterogeneous data formats, inconsistent variable naming, and inadequate quality control, which have hindered reproducible research across multiple disciplines. We present a high-quality, hourly-resolution meteorological dataset spanning 2000–2025 from 616 stations, featuring an innovative pipeline that automatically parses and semantically aligns heterogeneous Portuguese-language source data. A novel diagnostic quality control framework is introduced, preserving original values while applying two-stage checks for physical plausibility and spatiotemporal consistency. The resulting dataset includes standardized variables, unified timestamps, comprehensive metadata, and supporting audit files—such as station inventories, daily precipitation summaries, and variable-level failure statistics—significantly enhancing data transparency and usability for climate, environmental, agricultural, and machine learning applications.

Brazildata harmonizationmeteorological data

This study addresses the challenges of directly applying Global Climate Model (GCM) precipitation outputs to regional climate applications due to their non-Gaussianity, intermittency, and nonlinear representation of extremes. Conventional statistical and black-box machine learning approaches often lack interpretability, generalizability, and fidelity in preserving long-term trends. To overcome these limitations, this work proposes δCLIMBA (dCLIMBA)—the first differentiable modeling framework for GCM precipitation bias correction. dCLIMBA learns a physically informed, spatiotemporally adaptive mapping between CMIP6 historical simulations and Livneh reanalysis data through end-to-end training. The method accurately reproduces precipitation magnitudes and extreme-event quantile structures across multiple U.S. cities, achieves spatial performance comparable to LOCA2, preserves future climate trends, and substantially mitigates edge biases in unseen regions, offering a balanced combination of interpretability, generalization, and physical consistency.

bias correctionextreme eventsGlobal Circulation Model

This study addresses the limited physical interpretability of embedding representations derived from high-dimensional meteorological and climate data, which hinders similarity-based scientific discovery. To bridge this gap, the authors present an open-source visual analytics platform that tightly integrates embedding space exploration with meteorological physical context. By synergistically combining source data, metadata, and an external-memory-efficient retrieval mechanism, the platform enables interpretable and traceable analogy-driven discovery workflows. It overcomes conventional memory constraints, allowing efficient querying of extremely large embedding libraries on standard workstations. The approach is validated through a tropical cyclone event retrieval task: leveraging ERA5-derived embeddings and IBTrACS metadata, the system successfully performs analogical retrieval from known phenomena to previously unseen datasets.

embedding-based retrievallatent space interpretabilitymeteorological analogs

Hot Scholars

TA

Troy Arcomano

Allen Institute for AI
Machine learningAtmospheric Science
NB

Niklas Boers

Technical University of Munich, Potsdam Institute for Climate Impact Research, University of Exeter
Earth system dynamicsdata-driven modellingabrupt transitionsextreme events
LZ

Laure Zanna

New York University
Machine LearningPhysical OceanographyApplied MathematicsNumerical Modeling
AA

Alistair Adcroft

Princeton University
Numerical methodsOcean modelingIceberg modeling
OW

Oliver Watt-Meyer

Lead Research Scientist, Climate Modeling, Ai2
Atmospheric Dynamics and Climate