exponential family theory

Applying the theoretical structure of exponential-family distributions to derive factorization, sufficient-statistic representations, and closed-form expressions relevant to inference and in-context learning. This includes translating marginal constraints into log-linear factor graphs and analyzing resulting simplifications.

exponentialfamilytheory

12-Month Skill Trend

Momentum and market value over time
Trending
Score
+20 in 12 mo
96
12 mo agoNow
Career
Value
+$12K in 12 mo
$42K/year
12 mo agoNow

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

A Unified Theory of Exact Inference and Learning in Exponential Family Latent Variable Models

Apr 30, 2024
SS
Sacha Sokoloski
🏛️ Hertie Institute for AI in Brain Health | University of Tübingen

This work addresses the tractability of exact inference and learning in exponential-family latent variable models (LVMs), seeking to characterize the precise boundary of models admitting closed-form analytical solutions without approximation. Method: We derive necessary and sufficient conditions for prior–posterior conjugacy in exponential-family LVMs, providing the first systematic characterization of exact solvability. We further propose a composable graphical model construction framework that preserves structural flexibility while guaranteeing analytic tractability throughout. A general-purpose exact Bayesian inference and parameter learning algorithm is developed, accompanied by an open-source implementation supporting empirical validation across diverse models. Contribution/Results: Our results substantially broaden the class of LVMs amenable to exact inference—bypassing variational approximations or Monte Carlo sampling—and establish a rigorous theoretical foundation and practical toolkit for interpretable, high-precision latent-variable modeling.

Deriving necessary parameter constraints for tractable posterior distributionsDeveloping unified exact algorithms for inference and learningIdentifying exact inference conditions in exponential family latent variable models

Computational Approaches for Exponential-Family Factor Analysis

Mar 22, 2024
LW
Liang Wang
🏛️ Boston University

Existing exponential family factor analysis (EFFA) frameworks for non-Gaussian, missing, and heteroscedastic matrix data suffer from restrictive distributional assumptions and asymptotic bias in simulation-based maximum likelihood (SML) estimation. Method: We propose the first quasi-likelihood-based EFFA model, explicitly incorporating dispersion parameters and element-wise weights to enhance robustness against heteroscedasticity and arbitrary missingness mechanisms. We further design an EM-SGD hybrid algorithm that eliminates SML’s asymptotic bias, achieving a theoretical error bound of O(1/p) and enabling scalable inference. Results: Extensive experiments on synthetic data and three real-world modalities—count, binary, and skewed continuous matrices—demonstrate substantial improvements in low-rank covariance structure recovery and missing value imputation accuracy over state-of-the-art baselines.

Develops robust factor analysis for exponential-family data with quasi-likelihoodIntroduces dispersion and weights to handle large variations and missing valuesProvides efficient, unbiased computational methods scalable for large matrices

E-Values for Exponential Families: the General Case

Sep 17, 2024
YH
Yunda Hao
🏛️ the Chinese University of Hong Kong | CWI | Leiden University

This paper investigates the construction of e-variables and e-processes under composite exponential family null hypotheses. It systematically compares four approaches: reverse information projection (RIPr), conditional likelihood ratio (COND), universal inference (UI), and sequential RIPr. The work establishes, for the first time, the exact form of the RIPr prior in the Gaussian case and derives necessary and sufficient conditions for equivalence between RIPr and COND e-variables. Theoretically, it reveals a $(d/2)log n$ efficiency loss for UI and rigorously proves that COND is optimal in e-power. Precise expressions for e-power are derived for Gaussian models, and $o(1)$-accurate approximations are provided for general exponential families. The core contribution is a unifying framework that clarifies relationships among these methods and establishes COND as both theoretically optimal and practically implementable for e-variable construction.

Analyzing e-variables for composite exponential familiesCharacterizing RIPr prior for Gaussian and general alternativesComparing e-power of four e-statistics across sample sizes

This work proposes Exponential Family Discriminant Analysis (EFDA), a generalization of classical Linear Discriminant Analysis (LDA) that overcomes its reliance on Gaussian assumptions by accommodating any distribution within the exponential family. Under the assumption that class-conditional densities belong to the same exponential family, EFDA constructs linear decision rules based on sufficient statistics. This approach provides the first unified extension of the LDA framework beyond Gaussianity, guaranteeing asymptotic calibration and statistical efficiency while revealing structural calibration errors arising from model misspecification. Theoretical analysis integrates maximum likelihood estimation, the Cramér–Rao bound, and formal verification in Lean 4. Empirical evaluations across five non-Gaussian simulation settings demonstrate that EFDA matches or exceeds the classification accuracy of LDA, Quadratic Discriminant Analysis (QDA), and logistic regression, reduces expected calibration error by 2–6 times, and is the only method whose mean squared error consistently converges to zero.

CalibrationDiscriminant AnalysisExponential Family

Extremal graphical modeling with latent variables via convex optimization

Mar 14, 2024
SE
Sebastian Engelke
🏛️ University of Geneva | University of Washington

In multivariate extreme-value modeling, latent variables are pervasive and render conventional graphical models invalid. To address this, we propose eglent, the first tractable convex optimization method for learning Hüsler–Reiss extremal graphical models with latent variables. eglent integrates sparse–low-rank matrix decomposition into the extremal graph learning framework, jointly estimating the conditional dependence graph structure and the influence of latent variables—without requiring full observability. It consistently recovers both the graph structure and the number of latent variables. Theoretically, we establish finite-sample statistical consistency guarantees. Empirically, eglent significantly improves graph recovery accuracy on both synthetic and real-world datasets. Our core contribution is breaking the identifiability barrier in latent-variable extremal graph learning by introducing the first convex optimization paradigm that is computationally feasible, statistically interpretable, and theoretically grounded.

Decomposing precision matrix into sparse and low-rank componentsLearning extremal graphical models with latent variablesRecovering conditional graph and number of latent variables

Latest Papers

What's happening recently
View more

This work proposes a unified framework that systematically derives several classical results in exponential families through a concise identity involving the difference of Kullback–Leibler (KL) divergences and its inherent non-negativity. Relying solely on fundamental properties of KL divergence and the algebraic structure of exponential families, the approach reconstructs key results—such as the three-point and multi-point identities, the Pythagorean theorem in information geometry, and the Gibbs variational principle—without resorting to ad hoc or cumbersome proofs. Moreover, the framework naturally yields essential properties including the gradient formula for the log-partition function, the Bregman divergence representation, and the surjectivity of the moment map. These findings underscore the pivotal role of KL divergence as a unifying bridge linking information geometry, convex duality, and variational inference.

entropy-regularized reinforcement learningexponential familiesKL divergence

This work addresses the lack of intuition in traditional derivations of exponential family distributions, which often obscure their information-theoretic and physical foundations in pedagogical contexts. By leveraging the principle of maximum entropy and requiring only elementary notions of entropy, the paper presents a concise and self-contained derivation that avoids complex constrained optimization. The core contribution demonstrates that, under constraints fixing the expected values of sufficient statistics, exponential family distributions uniquely maximize relative entropy with respect to a general base measure, and Shannon entropy in the special case of a uniform base measure. This approach reveals the fundamental connection between maximum entropy and exponential families from minimal assumptions, substantially streamlining the didactic exposition and fostering deeper integration of statistical theory with physical reasoning.

exponential familiesfirst principlesinformation entropy

This work addresses the lack of non-asymptotic sample complexity guarantees for learning high-dimensional continuous exponential family distributions, particularly in the challenging setting of unbounded support. By employing score matching to perform structure learning on polynomial-form exponential family models, the study establishes the first finite-sample error bounds for this class of models through a synthesis of high-dimensional probabilistic modeling and non-asymptotic statistical analysis. The results demonstrate that the required sample size scales polynomially with the ambient dimension, thereby providing the first rigorous sample complexity guarantee for learning high-dimensional exponential family distributions with unbounded support. This contribution fills a critical theoretical gap in the non-asymptotic understanding of continuous exponential families.

exponential familieshigh-dimensional statisticsnon-asymptotic bounds

This work addresses the high computational cost and lack of convergence rate guarantees associated with nonparametric maximum likelihood estimation (NPMLE) in exponential family mixture models. The authors propose a data-compression-based acceleration strategy that, for the first time, reduces the likelihood evaluation complexity of NPMLE to logarithmic order. They establish rigorous statistical theory for the resulting approximate estimator, demonstrating that the proposed method achieves near-parametric convergence rates for marginal density estimation while substantially lowering computational overhead.

computation efficiencyconvergence rateexponential family mixtures

This work addresses the lack of a solid theoretical foundation for dependency networks, whose model distribution is implicitly defined as the stationary distribution of pseudo-Gibbs sampling and lacks a closed-form expression. From the perspective of information geometry, the paper interprets each step of pseudo-Gibbs sampling as an m-projection onto the manifold of full conditionals, thereby reformulating structure and parameter learning as a decomposable optimization problem. The authors introduce the full-conditional divergence and a tight upper bound to characterize the location of the stationary distribution within the probability simplex, and prove that the learned model converges uniformly to the true distribution as the sample size tends to infinity. Both theoretical analysis and empirical experiments demonstrate that the proposed upper bound is effective in practice and that the model enjoys statistical consistency.

dependency networksinformation geometrypseudo-Gibbs sampling

Hot Scholars

AK

Andrey Kupavskii

Moscow Institute of Physics and Technology
combinatoricsdiscrete geometry
HS

Helton Saulo

Assistant Professor of Statistics, University of Brasilia
EconometricsStatistical Learning
AM

Alessio Mansutti

IMDEA Software Institute
Formal verificationlogicmodel checking