Score
Designs, implements, and evaluates probabilistic mixture representations and estimation pipelines that model data as finite or sparse combinations of component distributions (discrete or continuous, e.g., Gaussian), including parametric mixtures and mixture density networks that map input vectors or embeddings to mixture parameters. Builds algorithms and training objectives for specifying component families and weights, fitting and regularizing mixtures (MDL, sparsity), training with normalization-free or CDF-matching losses, performing inference and sampling, approximating smoothed score functions, and imposing structural or consistency constraints so that downstream quantities (integrals, derivatives, or other functionals of the density) can be computed or satisfy required properties.
This paper addresses the efficient discrete approximation of Gaussian mixture models (GMMs) under the Wasserstein distance, motivated by dual requirements of quantization accuracy and computational scalability in control and cyber-physical system verification. We propose an enhanced quantization framework that integrates sigma-point sampling with adaptive clustering, enabling robust handling of high-dimensional, large-scale, and degenerate GMMs. A rigorous upper bound on the Wasserstein approximation error is derived, and a modular interface is provided to support customizable approximation schemes. Experiments demonstrate that our method achieves sublinear error convergence while significantly reducing computational overhead—outperforming state-of-the-art approaches in accuracy. The core contribution lies in the first systematic integration of sigma-point mechanisms with Wasserstein quantization theory, thereby unifying theoretical guarantees with practical deployability.
This work addresses key challenges in mixture distribution modeling—namely, the difficulty of integrating parametric and nonparametric approaches, poor compatibility across diverse distribution families (e.g., Gaussian, Poisson, kernel density), and limited scalability to high dimensions. To this end, we propose PMODE, a theoretically grounded, modular framework for mixture density estimation. PMODE partitions data into blocks and performs localized density estimation, seamlessly combining parametric and nonparametric components from heterogeneous distribution families while guaranteeing near-optimal convergence rates. Building upon this, we introduce MV-PMODE, an extension enabling scalable application to thousands of dimensions. Empirically, on the CIFAR-10 anomaly detection task, MV-PMODE achieves performance competitive with state-of-the-art deep generative models. This demonstrates a unified advance in modeling expressivity, theoretical rigor, and high-dimensional scalability.
Gibbs sampling for Bayesian mixture models suffers from slow mixing in the marginal posterior over component assignments and struggles to jointly perform model selection and parameter inference. Method: We propose two novel joint-sampling MCMC algorithms: (1) a collapsed Gibbs sampler incorporating unconventional move sets, and (2) a prior-driven, rejection-free component allocation sampler. Both methods jointly update observation assignments and the number of components, unifying model fitting and dimensionality inference. Contribution/Results: Our approaches eliminate the need for post-hoc model selection and substantially improve Markov chain mixing efficiency. In latent class analysis tasks, they reduce mixing time by several-fold compared to state-of-the-art methods while achieving comparable or superior posterior inference accuracy. The framework provides an efficient, fully automated computational solution for high-dimensional Bayesian nonparametric modeling.
This work addresses the problem of learning $k$-component Gaussian mixture models (GMMs) supported on the union of $k$ constant-radius balls in high dimensions, as well as Gaussian convolutions over low-dimensional manifolds or sets with small covering numbers. We propose the first analytically tractable diffusion-based algorithm for this setting. Methodologically, we unify score function estimation with higher-order Gaussian noise sensitivity analysis, augmented by poly-logarithmic piecewise polynomial regression and rigorous convergence theory. Under a minimal weight separation assumption, our algorithm achieves total variation error $varepsilon$ in quasi-polynomial time and sample complexity $Oig(n^{mathrm{poly}log((n+k)/varepsilon)}ig)$, overcoming fundamental limitations of classical algebraic approaches. Notably, this is the first subexponential learnability guarantee for Gaussian convolutions on manifolds. Moreover, our framework unifies and enables efficient learning of both continuous and discrete GMMs—resolving longstanding challenges in high-dimensional distribution learning.
Gaussian processes (GPs) suffer from cubic time complexity $O(N^3)$ and quadratic memory cost $O(N^2)$, limiting scalability to large-scale or nonstationary data. To address this, we propose the Mixture of Gaussian Process Experts (MoE-GP) model, capable of capturing nonstationarity, heteroscedasticity, and discontinuities. We introduce the first nested sequential Monte Carlo (SMC²) inference framework for joint Bayesian inference over both the gating network and GP expert parameters. Our approach preserves full parallelizability while significantly improving posterior estimation accuracy and stability—reducing variance compared to standard importance sampling. Experiments demonstrate strong robustness and scalability on complex temporal and spatial datasets where conventional stationary GPs fail. MoE-GP establishes a novel paradigm for scalable, nonstationary GP modeling.
To address the challenge of nonparametrically modeling mixed component distributions in heterogeneous data, this paper proposes a finite mixture model with nonparametric components, where each component density is itself modeled via a Dirichlet process mixture (DPM) prior. First, we establish identifiability conditions for the mixing components under this framework. Second, we theoretically prove that the posterior contraction rate for component densities is polynomial—significantly faster than the logarithmic rate typical of conventional deconvolution for mixing measures. Third, to enable efficient Bayesian inference, we design a tailored MCMC algorithm. Extensive simulations and real-data analyses demonstrate the method’s high accuracy and robustness in identifying latent subgroups, estimating population-level and component-specific densities. The approach thus offers both rigorous theoretical guarantees and practical utility for complex heterogeneous data analysis.
Modeling and equivalence reasoning for discrete-continuous hybrid probabilistic models—such as conditional Gaussian mixture models (CGMMs), where continuous variables follow multivariate Gaussians conditioned on discrete variables—remains challenging due to the lack of compositional, syntactic, and semantic foundations. Method: We introduce the first complete string diagram calculus for CGMMs, integrating categorical probability theory and compositional semantics to yield a graphical syntax with rigorous denotational meaning, accompanied by a sound and complete equational theory. Contribution/Results: This is the first framework to algebraically compose, visually represent, and precisely decide structural equivalence for such models: two diagrammatic expressions are equivalent if and only if they induce identical probability distributions. The calculus supports model construction, decomposition, optimization, and formal verification, thereby establishing a novel formal foundation for probabilistic programming and causal modeling.
Traditional neural networks suffer from rigid nonlinear expressivity due to fixed, hand-crafted activation functions (e.g., ReLU, Softmax). To address this, we propose the Gaussian Mixture Nonlinear Module (GMNM), which replaces static activations with a learnable, differentiable projection onto Gaussian kernels—thereby recasting nonlinearity modeling as a density approximation problem in a metric space. Inspired by Gaussian Mixture Models (GMMs), GMNM incorporates relaxed probabilistic constraints and a parameterized projection mechanism, enabling end-to-end training while maintaining architectural agnosticism across MLPs, CNNs, and Transformers. Extensive experiments on benchmark tasks—including image classification and language modeling—demonstrate consistent improvements in both accuracy and convergence speed. These results validate GMNM’s generality, effectiveness, and computational efficiency, offering a flexible, gradient-based alternative to conventional activation design.
This work addresses the slow convergence and mode collapse commonly encountered in traditional mixture density networks under maximum likelihood training. By reframing these networks as deep latent variable models, the study integrates the Expectation-Maximization (EM) framework with information geometry theory to propose, for the first time, a natural gradient EM (nGEM) objective function. This formulation reveals an intrinsic connection between mixture density networks and natural gradient descent. The resulting method achieves substantial improvements in training efficiency—accelerating convergence by up to tenfold—while maintaining robust performance on high-dimensional data, all with negligible additional computational overhead. Moreover, it effectively overcomes the failure modes associated with conventional negative log-likelihood optimization.
This work addresses the challenges in federated learning where the number of global clusters is unknown, local clustering structures are heterogeneous across clients, and clusters may overlap. To tackle these issues, the authors propose FedGEM, a novel federated clustering algorithm based on the Generalized Expectation-Maximization (GEM) framework. In FedGEM, each client performs local EM steps and constructs a component-wise uncertainty set, which the server leverages to infer the global number of clusters and identify cross-client cluster overlaps via a closed-form solution. As the first federated clustering method that jointly handles unknown global cluster counts and heterogeneous overlapping structures, FedGEM achieves low-complexity local computation and efficient aggregation. Theoretical analysis establishes its probabilistic convergence, and experiments demonstrate that its performance closely matches that of centralized EM while significantly outperforming existing federated clustering approaches.