Score
Designs and implements iterative parameter-estimation algorithms for incomplete-data or latent-variable models that alternate between computing expectations of hidden variables (E step) and performing maximization or conditional‑maximization updates (M/CM steps). This includes building EM and ECM variants to stabilize and accelerate likelihood maximization—e.g., separate updates for latent-factor and mixture parameters and scalable update rules for high‑dimensional data.
Standard Expectation-Maximization (EM) algorithms are ill-suited for statistical models with discrete latent variables due to their reliance on differentiability and continuous latent spaces. Method: We propose the first unified framework integrating Mirror Descent (MD) with Sequential Monte Carlo (SMC) for joint parameter estimation and posterior inference. Our approach jointly minimizes a variational functional over both the parameter space and the space of probability measures, enabling maximum likelihood estimation (MLE) without requiring latent variable continuity. Contribution/Results: This work breaks EM’s dependence on latent variable continuity, marking the first application of MD to MLE in discrete latent variable models, with rigorous convergence guarantees established. Experiments demonstrate significant improvements over standard EM across multiple discrete latent variable tasks; on real-valued latent variable benchmarks, our method matches state-of-the-art performance, validating both theoretical soundness and empirical robustness.
To address missing value imputation under incomplete training data, this paper proposes DiffPuter—a novel diffusion-based imputation framework. DiffPuter uniquely bridges unconditional diffusion training and conditional sampling with the M-step (learning the full-variable joint distribution) and E-step (iteratively imputing missing entries via conditional expectations) of the EM algorithm, enabling end-to-end trainable progressive imputation. It requires no pretraining, auxiliary discriminators, or assumptions about missingness mechanisms, and uniformly supports joint distribution modeling under arbitrary missing data patterns. Theoretically, we establish rigorous consistency between diffusion modeling and the EM paradigm. Empirically, DiffPuter achieves state-of-the-art performance across 10 benchmark datasets, outperforming 16 baselines with average reductions of 8.10% in MAE and 5.64% in RMSE—particularly excelling in complex, non-ignorable missingness scenarios.
This paper addresses inference for partially identified parameters in incomplete models. We propose a unified inferential method that simultaneously achieves robustness to model misspecification and information efficiency. Our core innovation is the first construction of a Kullback–Leibler (KL) information criterion that jointly accommodates both incompleteness and misspecification robustness, yielding a nonempty, identifiable set of pseudo-true parameters. The method fully exploits information from both discrete and continuous covariates and enables computationally tractable inference via an asymptotically pivotal Rao score statistic. We establish theoretical consistency and asymptotic normality under both correct specification and misspecification. Compared to existing approaches, our framework substantially enhances the reliability and applicability of partial identification inference, providing the first unified inferential framework for incomplete models with set-valued predictions that is both theoretically rigorous and practically implementable.
To address the challenge of scarce clean training data in Bayesian inverse problems, this paper proposes the first framework for learning theoretically grounded diffusion priors solely from incomplete and noisy observations. Methodologically, we embed diffusion probabilistic modeling into an Expectation-Maximization (EM) algorithm: the E-step estimates the latent variable posterior via iterative denoising sampling, while the M-step updates diffusion model parameters by maximizing the marginal likelihood. Our key contributions are: (1) the first provably consistent learning of diffusion priors directly from noisy and/or missing observations; and (2) an unconditional posterior sampling strategy that eliminates reliance on assumptions about the forward process. Experiments demonstrate that the learned prior achieves performance on par with fully supervised models in downstream inverse tasks—including denoising and inpainting—while ensuring rigorous generative consistency and theoretical guarantees.
In large-scale Gaussian process (GP) hyperparameter optimization, iterative linear solvers—such as conjugate gradient (CG)—induce inefficiency in computing gradients of the marginal likelihood due to repeated, costly matrix-vector operations. Method: We propose a general-purpose optimization framework integrating pathwise gradient estimation, solver warm-starting, and budget-aware early stopping. The framework is agnostic to the underlying iterative solver and supports CG, alternating projections, and stochastic gradient descent. Contribution/Results: Our approach substantially alleviates the accuracy–efficiency trade-off in gradient estimation. Experiments demonstrate up to 72× speedup over standard CG when solving to full convergence. Under early stopping, the average residual norm drops to one-seventh of that achieved by baseline methods, significantly shortening hyperparameter optimization time while preserving convergence stability and gradient estimation accuracy.
This study clarifies the fundamental distinction between restricted maximum likelihood (REML) and maximum likelihood (ML) estimation in linear mixed models. Within the EM algorithm framework, the two methods differ solely in their treatment of the covariance matrix during variance component updates: REML employs the prediction error covariance—corresponding to Henderson’s C matrix—whereas ML uses the conditional covariance. This work is the first to explicitly interpret REML’s computational rationale through the lens of prediction error covariance and provides concise R code to transparently illustrate the key matrices involved. The implemented algorithm successfully reproduces both ML and REML results from the lme4 package, clearly exposing the core difference in their covariance structures and offering a reproducible, pedagogically valuable tool for understanding and teaching these estimation methods.
This work addresses the limitations of the traditional Expectation–Maximization (EM) algorithm, which relies on ad hoc latent variable constructions and is restricted to specific missing-data settings, lacking a unified framework. The authors propose a Normalized EM (N-EM) algorithm that generalizes EM to log-likelihood optimization problems involving integral terms by introducing a normalized density function. This approach establishes a three-stage iterative scheme comprising a Normalization step (N-step), an Expectation step (E-step), and a Maximization step (M-step). For the first time, it provides a unified optimization framework applicable to a broader class of likelihood functions, eliminating the need for manually specified latent variables. The method not only solves problems intractable to conventional EM but also achieves efficient and consistent optimization in comparable scenarios. Theoretical analysis and extensive experiments confirm the convergence and effectiveness of the proposed algorithm.
本文提出了一种高效的EM算法,用于处理矩阵正态混合模型中的元素级和结构缺失问题,通过坐标近似更新条件均值和协方差,显著降低了计算成本。
研究高维数据缺失下的参数估计问题,通过统计与计算复杂性分析,揭示均值和协方差估计存在统计-计算差距,而线性回归则可通过高效算法接近信息论下界。
本文提出Impute-EM方法,通过交替填补缺失值和重新拟合扩散模型来处理异构数据中的缺失值问题,直接建模混合状态而不使用连续替代。