🤖 AI Summary
This study identifies a critical disconnect between generalizability and interpretability in machine learning–based agricultural yield forecasting: while models—including XGBoost, Random Forest, LSTM, and TCN—achieve high accuracy under spatially partitioned test sets at the German NUTS-3 level, their performance deteriorates substantially when evaluated on temporally independent validation years. Crucially, even under temporal generalization failure, SHAP-based attributions retain high confidence scores, revealing a fundamental limitation of post-hoc interpretability methods—the illusion of reliability. To address this, we propose “validation-aware interpretability,” a novel paradigm that grounds model explanation in robust spatiotemporal cross-validation. We rigorously demonstrate that high test-set accuracy does not imply reliable generalization, and that trustworthy interpretation must be predicated on temporal robustness. This work establishes validation-aware interpretability as a necessary condition for credible, actionable insights in time-sensitive agroecological modeling.
📝 Abstract
This study examines the generalization performance and interpretability of machine learning (ML) models used for predicting crop yield and yield anomalies in Germany's NUTS-3 regions. Using a high-quality, long-term dataset, the study systematically compares the evaluation and temporal validation behavior of ensemble tree-based models (XGBoost, Random Forest) and deep learning approaches (LSTM, TCN).
While all models perform well on spatially split, conventional test sets, their performance degrades substantially on temporally independent validation years, revealing persistent limitations in generalization. Notably, models with strong test-set accuracy, but weak temporal validation performance can still produce seemingly credible SHAP feature importance values. This exposes a critical vulnerability in post hoc explainability methods: interpretability may appear reliable even when the underlying model fails to generalize.
These findings underscore the need for validation-aware interpretation of ML predictions in agricultural and environmental systems. Feature importance should not be accepted at face value unless models are explicitly shown to generalize to unseen temporal and spatial conditions. The study advocates for domain-aware validation, hybrid modeling strategies, and more rigorous scrutiny of explainability methods in data-driven agriculture. Ultimately, this work addresses a growing challenge in environmental data science: how can we evaluate generalization robustly enough to trust model explanations?