Score
Designs, builds, and analyzes probabilistic models that are members of or approximated by exponential families by deriving their sufficient statistics, natural parameters, and factorization structure. This includes formulating marginal constraints as expectation constraints, deriving log-linear factor graphs and factorizations of joint distributions, and computing entropy-based lifts or dual representations for inference and learning.
This work addresses the tractability of exact inference and learning in exponential-family latent variable models (LVMs), seeking to characterize the precise boundary of models admitting closed-form analytical solutions without approximation. Method: We derive necessary and sufficient conditions for prior–posterior conjugacy in exponential-family LVMs, providing the first systematic characterization of exact solvability. We further propose a composable graphical model construction framework that preserves structural flexibility while guaranteeing analytic tractability throughout. A general-purpose exact Bayesian inference and parameter learning algorithm is developed, accompanied by an open-source implementation supporting empirical validation across diverse models. Contribution/Results: Our results substantially broaden the class of LVMs amenable to exact inference—bypassing variational approximations or Monte Carlo sampling—and establish a rigorous theoretical foundation and practical toolkit for interpretable, high-precision latent-variable modeling.
This work addresses the lack of intuition in traditional derivations of exponential family distributions, which often obscure their information-theoretic and physical foundations in pedagogical contexts. By leveraging the principle of maximum entropy and requiring only elementary notions of entropy, the paper presents a concise and self-contained derivation that avoids complex constrained optimization. The core contribution demonstrates that, under constraints fixing the expected values of sufficient statistics, exponential family distributions uniquely maximize relative entropy with respect to a general base measure, and Shannon entropy in the special case of a uniform base measure. This approach reveals the fundamental connection between maximum entropy and exponential families from minimal assumptions, substantially streamlining the didactic exposition and fostering deeper integration of statistical theory with physical reasoning.
This paper addresses the challenge of modeling uncertainty by systematically establishing a pedagogical and theoretical framework for probabilistic graphical models (PGMs). To tackle the intractability of representing and reasoning over high-dimensional joint distributions, it unifies directed graphs (Bayesian networks) and undirected graphs (Markov random fields) to compactly encode variable dependencies, integrating probability theory with graph theory. The work develops a comprehensive methodology encompassing parameter learning, structure learning, and exact/approximate inference—including variable elimination, belief propagation, and variational inference. Its primary contribution is a tripartite PGM pedagogical paradigm—representation, learning, and inference—that rigorously aligns graph structure with probabilistic semantics. Through algorithmic design and concrete case studies, the framework enhances model interpretability and practical utility in prediction and decision-making tasks, thereby providing foundational support for uncertainty reasoning in machine learning and AI.
Existing exponential family factor analysis (EFFA) frameworks for non-Gaussian, missing, and heteroscedastic matrix data suffer from restrictive distributional assumptions and asymptotic bias in simulation-based maximum likelihood (SML) estimation. Method: We propose the first quasi-likelihood-based EFFA model, explicitly incorporating dispersion parameters and element-wise weights to enhance robustness against heteroscedasticity and arbitrary missingness mechanisms. We further design an EM-SGD hybrid algorithm that eliminates SML’s asymptotic bias, achieving a theoretical error bound of O(1/p) and enabling scalable inference. Results: Extensive experiments on synthetic data and three real-world modalities—count, binary, and skewed continuous matrices—demonstrate substantial improvements in low-rank covariance structure recovery and missing value imputation accuracy over state-of-the-art baselines.
This paper investigates the construction of e-variables and e-processes under composite exponential family null hypotheses. It systematically compares four approaches: reverse information projection (RIPr), conditional likelihood ratio (COND), universal inference (UI), and sequential RIPr. The work establishes, for the first time, the exact form of the RIPr prior in the Gaussian case and derives necessary and sufficient conditions for equivalence between RIPr and COND e-variables. Theoretically, it reveals a $(d/2)log n$ efficiency loss for UI and rigorously proves that COND is optimal in e-power. Precise expressions for e-power are derived for Gaussian models, and $o(1)$-accurate approximations are provided for general exponential families. The core contribution is a unifying framework that clarifies relationships among these methods and establishes COND as both theoretically optimal and practically implementable for e-variable construction.
This work investigates how to learn interpretable abstract rules from probabilistic signals and establishes their connection to probabilistic graphical models. It proposes a learning framework grounded in information lattices, where rules are interpreted as marginal constraints on quotient variables by alternately projecting signals onto partition lattices and lifting the resulting rules back to the original domain. This approach explicitly links information lattice learning to constraint-based factor graph learning for the first time, revealing an intrinsic relationship with maximum entropy models and introducing a novel interpretable modeling paradigm centered on quotient variables. The contributions include deriving the corresponding log-linear factor graph representation, clarifying the fundamental distinction between information lattices and Bayesian networks, and opening new avenues for hybrid symbolic–probabilistic learning.
This work addresses the high computational cost and lack of convergence rate guarantees associated with nonparametric maximum likelihood estimation (NPMLE) in exponential family mixture models. The authors propose a data-compression-based acceleration strategy that, for the first time, reduces the likelihood evaluation complexity of NPMLE to logarithmic order. They establish rigorous statistical theory for the resulting approximate estimator, demonstrating that the proposed method achieves near-parametric convergence rates for marginal density estimation while substantially lowering computational overhead.
This work proposes a unified framework that systematically derives several classical results in exponential families through a concise identity involving the difference of Kullback–Leibler (KL) divergences and its inherent non-negativity. Relying solely on fundamental properties of KL divergence and the algebraic structure of exponential families, the approach reconstructs key results—such as the three-point and multi-point identities, the Pythagorean theorem in information geometry, and the Gibbs variational principle—without resorting to ad hoc or cumbersome proofs. Moreover, the framework naturally yields essential properties including the gradient formula for the log-partition function, the Bregman divergence representation, and the surjectivity of the moment map. These findings underscore the pivotal role of KL divergence as a unifying bridge linking information geometry, convex duality, and variational inference.
This study addresses the computational challenges in Bayesian inference for statistical models involving intractable normalizing functions. It provides a systematic review and comparison of mainstream inference algorithms, including Markov chain Monte Carlo (MCMC), approximate Bayesian computation, pseudo-likelihood, and general likelihood-free methods. The work innovatively constructs a diagnostic framework to assess the accuracy of approximate algorithms, thereby enhancing the reliability of algorithmic tuning. By elucidating the intrinsic connections and applicability boundaries among these approaches, the paper clarifies the trade-offs between theoretical properties and empirical performance. Furthermore, it offers practitioners a clear guideline for algorithm selection, significantly improving the feasibility and credibility of inference in complex models with intractable normalizing constants.
This study elucidates the implicit informational assumptions embedded in Bayesian hierarchical models, with a focus on the nature of constraints arising from dependencies among parameters. By integrating the principle of maximum entropy, probability integral transforms, and marginalization analysis, the work rigorously demonstrates that when hyperpriors are specified as maximum entropy distributions, the induced marginal priors retain the maximum entropy structure, with constraints acting directly on the marginal distributions of functions of unknown quantities. This result clarifies the informational content encoded in hierarchical priors, deepens the understanding of their semantic meaning and structural constraints, and provides theoretical foundations for enhancing the interpretability and principled design of Bayesian models.