LimeSoDa: A Dataset Collection for Benchmarking of Machine Learning Regressors in Digital Soil Mapping

📅 2025-02-27
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Evaluating statistical methods in digital soil mapping (DSM) is hindered by closed, single-source datasets, limiting the generalizability of conclusions. Method: We introduce LimeSoDa—the first open, multi-source, standardized benchmark dataset for DSM—comprising 31 field- or farm-scale datasets from diverse countries, uniformly providing three target variables (soil organic matter, clay content, pH) and spectral/sensing-derived features. We systematically harmonize heterogeneous soil sensing data into ready-to-use tabular formats and propose a context-aware evaluation framework integrating multiple algorithms and realistic scenarios. Contribution/Results: Benchmark experiments reveal that model performance critically depends on feature dimensionality and data provenance: MLR and SVR excel with high-dimensional spectral data, whereas CatBoost and RF outperform in low-dimensional settings (<20 features). LimeSoDa establishes a reproducible, extensible infrastructure for rigorous, comparative assessment of DSM methodologies.

Technology Category

Application Category

📝 Abstract
Digital soil mapping (DSM) relies on a broad pool of statistical methods, yet determining the optimal method for a given context remains challenging and contentious. Benchmarking studies on multiple datasets are needed to reveal strengths and limitations of commonly used methods. Existing DSM studies usually rely on a single dataset with restricted access, leading to incomplete and potentially misleading conclusions. To address these issues, we introduce an open-access dataset collection called Precision Liming Soil Datasets (LimeSoDa). LimeSoDa consists of 31 field- and farm-scale datasets from various countries. Each dataset has three target soil properties: (1) soil organic matter or soil organic carbon, (2) clay content and (3) pH, alongside a set of features. Features are dataset-specific and were obtained by optical spectroscopy, proximal- and remote soil sensing. All datasets were aligned to a tabular format and are ready-to-use for modeling. We demonstrated the use of LimeSoDa for benchmarking by comparing the predictive performance of four learning algorithms across all datasets. This comparison included multiple linear regression (MLR), support vector regression (SVR), categorical boosting (CatBoost) and random forest (RF). The results showed that although no single algorithm was universally superior, certain algorithms performed better in specific contexts. MLR and SVR performed better on high-dimensional spectral datasets, likely due to better compatibility with principal components. In contrast, CatBoost and RF exhibited considerably better performances when applied to datasets with a moderate number (<20) of features. These benchmarking results illustrate that the performance of a method is highly context-dependent. LimeSoDa therefore provides an important resource for improving the development and evaluation of statistical methods in DSM.
Problem

Research questions and friction points this paper is trying to address.

Benchmarking machine learning regressors in soil mapping.
Evaluating algorithm performance across diverse soil datasets.
Addressing dataset access limitations in digital soil mapping.
Innovation

Methods, ideas, or system contributions that make the work stand out.

Open-access dataset collection
Multiple learning algorithms comparison
Context-dependent performance analysis
🔎 Similar Papers
No similar papers found.
J
J. Schmidinger
Osnabrück University, Joint Lab Artificial Intelligence and Data Science, Osnabrück, Germany; Leibniz Institute for Agricultural Engineering and Bioeconomy (ATB), Department of Agromechatronics, Potsdam, Germany
Sebastian Vogel
Sebastian Vogel
NXP Semiconductors
Efficient Processing of CNNsDeep LearningComputer VisionEmbedded AI
Viacheslav Barkov
Viacheslav Barkov
Doctoral Research Assistant, Osnabrück University
Anh-Duy Pham
Anh-Duy Pham
Julius-Maximilians-Universität Würzburg
Operations ResearchGraph LearningAgent-based Modelling
R
R. Gebbers
Leibniz Institute for Agricultural Engineering and Bioeconomy (ATB), Department of Agromechatronics, Potsdam, Germany
Hamed Tavakoli
Hamed Tavakoli
Leibniz Institute for Agricultural Engineering and Bioeconomy
Precision agriculturesensors and sensor systems for plant and soil monitoringProximal soil
Jose Correa
Jose Correa
Universidad de Chile
Operations ResearchMarket DesignAlgorithmic Game TheoryMechanism Design
T
Tiago R. Tavares
University of São Paulo (USP), Center of Nuclear Energy in Agriculture (CENA), Piracicaba, Brazil
P
Patrick Filippi
The University of Sydney, Sydney Institute of Agriculture, Sydney, Australia
E
Edward J Jones
The University of Sydney, Sydney Institute of Agriculture, Sydney, Australia
V
Vojtech Lukas
Mendel University in Brno, Department of Agrosystems and Bioclimatology, Brno, Czech Republic
E
Eric Boenecke
Leibniz Institute of Vegetable and Ornamental Crops, Next Generation Horticultural Systems, Grossbeeren, Germany
J
J. Ruehlmann
Leibniz Institute of Vegetable and Ornamental Crops, Next Generation Horticultural Systems, Grossbeeren, Germany
I
Ingmar Schroeter
Eberswalde University for Sustainable Development, Landscape Management and Nature Conservation, Eberswalde, Germany
E
E. Kramer
Eberswalde University for Sustainable Development, Landscape Management and Nature Conservation, Eberswalde, Germany
S
Stefan Paetzold
University of Bonn, Institute of Crop Science and Resource Conservation (INRES) —Soil Science and Soil Ecology, Bonn, Germany
M
Masakazu Kodaira
Tokyo University of Agriculture and Technology, Institute of Agriculture, Tokyo, Japan
.
.J.-C. Wadoux
LISAH, Univ. Montpellier, AgroParisTech, INRAE, IRD, L'Institut Agro, Montpellier, France
L
Luca Bragazza
Agroscope, Field-Crop Systems and Plant Nutrition, Nyon, Switzerland
K
Konrad Metzger
Agroscope, Field-Crop Systems and Plant Nutrition, Nyon, Switzerland
J
Jingyi Huang
University of Wisconsin-Madison, Department of Soil Science, Madison, USA
D
D. S. M. Valente
Federal University of Viçosa, Department of Agricultural Engineering, Viçosa, Brazil
J
J. L. Safanelli
Woodwell Climate Research Center, Falmouth, USA
E
Eduardo L. Bottega
Federal University of Santa Maria (UFSM), Academic Coordination, Santa Maria, Brazil
R
R. S. Dalmolin
Federal University of Santa Maria (UFSM), Academic Coordination, Santa Maria, Brazil
Csilla Farkas
Csilla Farkas
A
Alexander Steiger
T
Taciara Zborowski Horst
L
L. Ramirez-Lopez
Thomas Scholten
Thomas Scholten
Professor of Soil Science and Geomorphology
Soil ScienceExtreme EnvironmentsGeomorphologyGeoecologySoil Erosion
F
Felix Stumpf
P
Pablo Rosso
M
Marcelo M. Costa
R
R. S. Zandonadi
J
Johanna Wetterlind
M
M. Atzmueller
Osnabrück University, Joint Lab Artificial Intelligence and Data Science, Osnabrück, Germany