🤖 AI Summary
This study addresses the limited model interpretability and insufficient cross-ecoregion information utilization in wildfire prediction across the western United States by proposing a multi-level Bayesian regression tree model. Methodologically, it integrates data from multiple ecoregions through shared hyperparameters and introduces a novel sampling algorithm based on parallel tempering to optimize conditional mixing at posterior temperatures, thereby overcoming the MCMC sampling bottleneck inherent in traditional Bayesian CART models. Results demonstrate that the proposed approach significantly enhances tree structure similarity and out-of-sample predictive performance. Furthermore, it precisely identifies potential evapotranspiration, temperature, and evergreen forest cover as critical hazard-inducing factors, achieving simultaneous improvements in both predictive accuracy and interpretability.
📝 Abstract
We propose a Bayesian regression tree model fit within a multilevel structure and apply it to historic wildfire data in the western United States. Sharing of information between related groups (ecoregions) combined with highly interpretable regression trees allows for better predictions and understanding of climate and land cover variables predictive of wildfires. By doing a simulation study with a range of performance metrics, we demonstrate our method produces tree posteriors most structurally similar to assumed true trees, while simultaneously achieving good out-of-sample predictive performance. Applied to a large wildfire data set, we explore variable splits within regression trees corresponding to each ecoregion in detail, taking into account known features of each location. Shared hyperparameters between trees provide highly useful understanding of both variable and split value importance in predicting wildfires among all ecoregions, with no direct parallel in comparable models. Namely, we highlight potential evaporation, temperature, and evergreen forest land cover as variables most associated with historic wildfires, with some observable patterns in split values most commonly chosen across the groups. We propose a new algorithm based on parallel tempering, conditioning on shared hyperparameters at the true posterior temperature, improving Markov chain mixing, a known bottleneck in Bayesian CART models.