METBRA25Y: Brazil Surface Meteorology Archive with Harmonized Variables and Quality Control

📅 2026-05-09
📈 Citations: 0
Influential: 0
📄 PDF

career value

184K/year
🤖 AI Summary
This study addresses long-standing challenges in Brazilian surface meteorological observations, including heterogeneous data formats, inconsistent variable naming, and inadequate quality control, which have hindered reproducible research across multiple disciplines. We present a high-quality, hourly-resolution meteorological dataset spanning 2000–2025 from 616 stations, featuring an innovative pipeline that automatically parses and semantically aligns heterogeneous Portuguese-language source data. A novel diagnostic quality control framework is introduced, preserving original values while applying two-stage checks for physical plausibility and spatiotemporal consistency. The resulting dataset includes standardized variables, unified timestamps, comprehensive metadata, and supporting audit files—such as station inventories, daily precipitation summaries, and variable-level failure statistics—significantly enhancing data transparency and usability for climate, environmental, agricultural, and machine learning applications.
📝 Abstract
This data paper describes METBRA25Y, a harmonized archive of hourly surface meteorological observations from Brazil derived from public historical records of the Instituto Nacional de Meteorologia (INMET). The dataset was designed to support reproducible environmental, climatological, hydrological, agricultural, urban-risk, and machine-learning studies that require station-level meteorological time series with standardized variable names and explicit quality-control metadata. The processing workflow ingests annual INMET archives, parses station metadata from raw file headers, normalizes heterogeneous Portuguese column names into a canonical schema, constructs hourly timestamps, consolidates observations by city and station, and exports compressed CSV files together with station manifests, per-station quality flags, daily precipitation aggregates, variable-level failure summaries, and missing-data audits. The quality-control protocol follows a two-stage strategy: first, physically implausible values are converted to missing values and flagged; second, temporal and cross-variable consistency checks generate diagnostic flags without necessarily overwriting the original measurements. The resulting package covers observations between 2000 and 2025, with stationspecific temporal coverage, and includes key meteorological variables such as precipitation, air temperature, dew point, relative humidity, atmospheric pressure, wind speed, wind gust, wind direction, and global solar radiation. Based on the summary files included in the current release snapshot, the archive contains 616 unique station codes across variable summaries, of which 605 have coordinates within a broad Brazil plausibility envelope. This paper documents the dataset provenance, file organization, harmonized schema, quality-control rules, technical validation outputs, limitations, and recommended usage practices.
Problem

Research questions and friction points this paper is trying to address.

meteorological data
data harmonization
quality control
surface observations
Brazil
Innovation

Methods, ideas, or system contributions that make the work stand out.

harmonized meteorological data
quality control protocol
two-stage QC strategy
standardized schema
missing-data audit
🔎 Similar Papers
No similar papers found.