expectation-conditional maximisation

Designs and implements iterative parameter-estimation algorithms for incomplete-data or latent-variable models that alternate between computing expectations of hidden variables (E step) and performing maximization or conditional‑maximization updates (M/CM steps). This includes building EM and ECM variants to stabilize and accelerate likelihood maximization—e.g., separate updates for latent-factor and mixture parameters and scalable update rules for high‑dimensional data.

expectation-conditionalmaximisation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.17
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

A mirror descent approach to maximum likelihood estimation in latent variable models

Jan 27, 2025
FR
F. R. Crucinio
🏛️ University of Turin | Collegio Carlo Alberto

Standard Expectation-Maximization (EM) algorithms are ill-suited for statistical models with discrete latent variables due to their reliance on differentiability and continuous latent spaces. Method: We propose the first unified framework integrating Mirror Descent (MD) with Sequential Monte Carlo (SMC) for joint parameter estimation and posterior inference. Our approach jointly minimizes a variational functional over both the parameter space and the space of probability measures, enabling maximum likelihood estimation (MLE) without requiring latent variable continuity. Contribution/Results: This work breaks EM’s dependence on latent variable continuity, marking the first application of MD to MLE in discrete latent variable models, with rigorous convergence guarantees established. Experiments demonstrate significant improvements over standard EM across multiple discrete latent variable tasks; on real-valued latent variable benchmarks, our method matches state-of-the-art performance, validating both theoretical soundness and empirical robustness.

Enables estimation when latent variables take discrete valuesOutperforms standard expectation maximisation algorithms in performancePerforms joint parameter inference and posterior estimation in latent variable models

Unleashing the Potential of Diffusion Models for Incomplete Data Imputation

May 31, 2024
HZ
Hengrui Zhang
🏛️ University of Illinois at Chicago

To address missing value imputation under incomplete training data, this paper proposes DiffPuter—a novel diffusion-based imputation framework. DiffPuter uniquely bridges unconditional diffusion training and conditional sampling with the M-step (learning the full-variable joint distribution) and E-step (iteratively imputing missing entries via conditional expectations) of the EM algorithm, enabling end-to-end trainable progressive imputation. It requires no pretraining, auxiliary discriminators, or assumptions about missingness mechanisms, and uniformly supports joint distribution modeling under arbitrary missing data patterns. Theoretically, we establish rigorous consistency between diffusion modeling and the EM paradigm. Empirically, DiffPuter achieves state-of-the-art performance across 10 benchmark datasets, outperforming 16 baselines with average reductions of 8.10% in MAE and 5.64% in RMSE—particularly excelling in complex, non-ignorable missingness scenarios.

Addressing missing data imputation using diffusion modelsEnhancing conditional sampling for accurate value updatesOvercoming training incompleteness in generative models

Information based inference in models with set-valued predictions and misspecification

Jan 19, 2024
HK
Hiroaki Kaido
🏛️ Boston University | Cornell University

This paper addresses inference for partially identified parameters in incomplete models. We propose a unified inferential method that simultaneously achieves robustness to model misspecification and information efficiency. Our core innovation is the first construction of a Kullback–Leibler (KL) information criterion that jointly accommodates both incompleteness and misspecification robustness, yielding a nonempty, identifiable set of pseudo-true parameters. The method fully exploits information from both discrete and continuous covariates and enables computationally tractable inference via an asymptotically pivotal Rao score statistic. We establish theoretical consistency and asymptotic normality under both correct specification and misspecification. Compared to existing approaches, our framework substantially enhances the reliability and applicability of partial identification inference, providing the first unified inferential framework for incomplete models with set-valued predictions that is both theoretically rigorous and practically implementable.

Develops inference for partially identified parameters in incomplete modelsHandles both correctly specified and misspecified model scenariosUses Kullback-Leibler criterion and Rao's score statistic for robust inference

Learning Diffusion Priors from Observations by Expectation Maximization

May 22, 2024
FR
François Rozet
🏛️ University of Liège | Université Paris-Saclay | Université Paris Cité | CEA | CNRS | AIM

To address the challenge of scarce clean training data in Bayesian inverse problems, this paper proposes the first framework for learning theoretically grounded diffusion priors solely from incomplete and noisy observations. Methodologically, we embed diffusion probabilistic modeling into an Expectation-Maximization (EM) algorithm: the E-step estimates the latent variable posterior via iterative denoising sampling, while the M-step updates diffusion model parameters by maximizing the marginal likelihood. Our key contributions are: (1) the first provably consistent learning of diffusion priors directly from noisy and/or missing observations; and (2) an unconditional posterior sampling strategy that eliminates reliance on assumptions about the forward process. Experiments demonstrate that the learned prior achieves performance on par with fully supervised models in downstream inverse tasks—including denoising and inpainting—while ensuring rigorous generative consistency and theoretical guarantees.

Improving posterior sampling for diffusion modelsLearning priors from noisy observations onlyTraining diffusion models without clean data

Improving Linear System Solvers for Hyperparameter Optimisation in Iterative Gaussian Processes

May 28, 2024
JA
Jihao Andreas Lin
🏛️ University of Cambridge | MPI for Intelligent Systems

In large-scale Gaussian process (GP) hyperparameter optimization, iterative linear solvers—such as conjugate gradient (CG)—induce inefficiency in computing gradients of the marginal likelihood due to repeated, costly matrix-vector operations. Method: We propose a general-purpose optimization framework integrating pathwise gradient estimation, solver warm-starting, and budget-aware early stopping. The framework is agnostic to the underlying iterative solver and supports CG, alternating projections, and stochastic gradient descent. Contribution/Results: Our approach substantially alleviates the accuracy–efficiency trade-off in gradient estimation. Experiments demonstrate up to 72× speedup over standard CG when solving to full convergence. Under early stopping, the average residual norm drops to one-seventh of that achieved by baseline methods, significantly shortening hyperparameter optimization time while preserving convergence stability and gradient estimation accuracy.

Gaussian ProcessesHyperparameter OptimizationLarge-scale Datasets

Latest Papers

What's happening recently
View more

This study clarifies the fundamental distinction between restricted maximum likelihood (REML) and maximum likelihood (ML) estimation in linear mixed models. Within the EM algorithm framework, the two methods differ solely in their treatment of the covariance matrix during variance component updates: REML employs the prediction error covariance—corresponding to Henderson’s C matrix—whereas ML uses the conditional covariance. This work is the first to explicitly interpret REML’s computational rationale through the lens of prediction error covariance and provides concise R code to transparently illustrate the key matrices involved. The implemented algorithm successfully reproduces both ML and REML results from the lme4 package, clearly exposing the core difference in their covariance structures and offering a reproducible, pedagogically valuable tool for understanding and teaching these estimation methods.

EM algorithmlinear mixed modelsprediction-error covariances

This work addresses the limitations of the traditional Expectation–Maximization (EM) algorithm, which relies on ad hoc latent variable constructions and is restricted to specific missing-data settings, lacking a unified framework. The authors propose a Normalized EM (N-EM) algorithm that generalizes EM to log-likelihood optimization problems involving integral terms by introducing a normalized density function. This approach establishes a three-stage iterative scheme comprising a Normalization step (N-step), an Expectation step (E-step), and a Maximization step (M-step). For the first time, it provides a unified optimization framework applicable to a broader class of likelihood functions, eliminating the need for manually specified latent variables. The method not only solves problems intractable to conventional EM but also achieves efficient and consistent optimization in comparable scenarios. Theoretical analysis and extensive experiments confirm the convergence and effectiveness of the proposed algorithm.

data augmentationexpectation-maximizationintegral

Hot Scholars

GX

Gongjun Xu

University of Michigan
StatisticsMachine LearningPsychometrics
AP

Antonio Punzo

Full Professor of Statistics, University of Catania
Mixture ModelsHidden Markov ModelsHeavy-tailed DistributionsSerial Dependence
MB

Michael Burke

Monash University
Robot learningImitation learningIntelligent RoboticsMachine Learning