Hierarchical random measures without tables

πŸ“… 2025-05-05
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
HDP-based hierarchical Bayesian modeling suffers from high posterior computational complexity and rapidly deteriorating sampling efficiency with increasing data size and hierarchy depth, primarily due to the reliance on latent β€œtable” variables. To address this, we propose a table-free hierarchical random measure modeling framework. Our contributions are threefold: (1) We introduce a novel prior for concentration parameters that induces quasi-conjugate posteriors, eliminating dependence on table variables entirely; (2) We generalize the approach to the broad class of normalized hierarchical random measures; (3) Leveraging multivariate incremental independence and completely random vector representations, we derive an efficient and exact Gibbs sampler family. Experiments demonstrate substantial gains in modeling efficiency on large-scale, deep-hierarchical datasets. The method preserves statistical consistency while yielding more interpretable posterior representations and significantly accelerated convergence.

Technology Category

Reasoning under Uncertainty: Relational Probabilistic ModelsMachine Learning: Probabilistic Circuits and Graphical ModelsKnowledge Representation and Reasoning: Computational Complexity of Reasoning

Application Category

Graph Algorithms and Modeling for the Web: Efficient manipulation of static and dynamic Web-related graphsUser Modeling, Personalization and Recommendation: Metrics for user behavior and evaluating successSecurity and Privacy: Large-scale security measurements
πŸ“ Abstract
The hierarchical Dirichlet process is the cornerstone of Bayesian nonparametric multilevel models. Its generative model can be described through a set of latent variables, commonly referred to as tables within the popular restaurant franchise metaphor. The latent tables simplify the expression of the posterior and allow for the implementation of a Gibbs sampling algorithm to approximately draw samples from it. However, managing their assignments can become computationally expensive, especially as the size of the dataset and of the number of levels increase. In this work, we identify a prior for the concentration parameter of the hierarchical Dirichlet process that (i) induces a quasi-conjugate posterior distribution, and (ii) removes the need of tables, bringing to more interpretable expressions for the posterior, with both a faster and an exact algorithm to sample from it. Remarkably, this construction extends beyond the Dirichlet process, leading to a new framework for defining normalized hierarchical random measures and a new class of algorithms to sample from their posteriors. The key analytical tool is the independence of multivariate increments, that is, their representation as completely random vectors.
Problem

Research questions and friction points this paper is trying to address.

Eliminates computational cost of managing latent tables in hierarchical models
Introduces quasi-conjugate posterior for hierarchical Dirichlet process
Extends framework to normalized hierarchical random measures
Innovation

Methods, ideas, or system contributions that make the work stand out.

Quasi-conjugate posterior for concentration parameter
Eliminates tables for faster computation
Extends to normalized hierarchical random measures
πŸ”Ž Similar Papers
No similar papers found.
M
Marta Catalano
Luiss University, Rome, Italy
C
Claudio Del Sole
University of Milan - Bicocca, Milan, Italy