🤖 AI Summary
This study addresses three key challenges in applying Bayesian finite mixture models to equipment degradation risk clustering—sparse signals, unstable clustering, and computational infeasibility of MCMC—by proposing an efficient and stable heterogeneous degradation risk clustering framework. The core innovations include the first empirical validation of the critical role of 8-state global percentile discretization in enhancing model stability; the construction of a 30-dimensional multi-source feature engineering pipeline integrating statistical, continuous, and semantic features (including PCA-compressed text embeddings); and the design of an interpretable, anti-overfitting model selection criterion based on WAIC, augmented with constraints on minimum cluster proportion and separation. Replacing NUTS with full-rank automatic differentiation variational inference (ADVI), the approach achieves 84× speedup over NUTS while maintaining stable results, shows 15× faster convergence with high estimation consistency under random-effects models, and accurately identifies the optimal interpretable risk clusters, as demonstrated on a dataset of 280 industrial pumps and 104,703 inspection records.
📝 Abstract
Bayesian finite mixture models can identify discrete risk clusters (low-risk vs. high-risk equipment), but face three critical bottlenecks: (1) insufficient degradation signals from coarse state discretization, (2) unstable cluster identification when data inherently supports fewer clusters than explored, and (3) computational infeasibility of Markov Chain Monte Carlo (MCMC) methods for production deployment (7+ hours per model). We propose a practical framework combining (1) 8-state global percentile discretization that amplifies degradation events, (2) 30-dimensional feature engineering integrating statistical trends (22 features), continuous health indicators, and text embeddings (PCA-compressed to 3 dimensions), (3) interpretable model selection rules enforcing minimum cluster share and separation alongside WAIC, and (4) Automatic Differentiation Variational Inference (ADVI) with full-rank covariance for stable, fast estimation. Applied to 280 industrial pump equipment with 104,703 inspection records, we demonstrate: (1) Random effect models (baseline) show ADVI and NUTS produce nearly identical estimates with 15$\times$ speedup, validating ADVI accuracy. (2) Finite mixture models identify optimal number of clusters with interpretability constraints. (3) NUTS exhibits severe convergence issues and label switching, while ADVI provides stable results in 84$\times$ less time. We contributed that (1) First demonstration that fine-grained state discretization (8-state) is essential for mixture model stability in survival analysis.(2) Comprehensive feature engineering strategy combining statistical, continuous, and semantic signals. (3) Practical interpretability rules preventing overfitting in automated model selection. (4) Empirical evidence that ADVI outperforms NUTS for finite mixture models in terms of convergence, stability, and computational efficiency.