🤖 AI Summary
To address particle degeneracy—caused by iterative resampling in recursive Bayesian inference—this paper proposes a novel method for constructing proposal distributions via smoothed empirical distributions. At each recursion step, the method leverages the empirical distribution of historical particles, applies kernel smoothing to yield a high-quality proposal, and then employs accept-reject sampling for efficient importance weighting and resampling, thereby balancing computational efficiency and particle diversity. Its key innovation lies in embedding empirical distribution smoothing directly into the recursive framework, mitigating rapid degeneration inherent in conventional sequential importance sampling (SIS) and particle filtering (PF). Empirical evaluation on simulated data from logistic regression and a hierarchical forest vegetation model for New Mexico demonstrates that the method significantly improves posterior approximation accuracy and stability. It is particularly effective for streaming data and large-scale block-wise analysis.
📝 Abstract
Recursive Bayesian inference, in which posterior beliefs are updated in light of accumulating data, is a tool for implementing Bayesian models in applications with streaming and/or very large data sets. As the posterior of one iteration becomes the prior for the next, beliefs are updated sequentially instead of all-at-once. Thus, recursive inference is relevant for both streaming data and settings where data too numerous to be analyzed together can be partitioned into manageable pieces. In practice, posteriors are characterized by samples obtained using, e.g., acceptance/rejection sampling in which draws from the posterior of one iteration are used as proposals for the next. While simple to implement, such filtering approaches suffer from particle depletion, degrading each sample's ability to represent its target posterior. As a remedy, we investigate generating proposals from a smoothed version of the preceding sample's empirical distribution. The method retains computationally valuable properties of similar methods, but without particle depletion, and we demonstrate its accuracy in simulation. We apply the method to data simulated from both a simple, logistic regression model as well as a hierarchical model originally developed for classifying forest vegetation in New Mexico using satellite imagery.