A Bayesian approach to modeling topic-metadata relationships

📅 2021-04-06
🏛️ AStA Advances in Statistical Analysis
📈 Citations: 3
✨ Influential: 1
📄 PDF
🤖 AI Summary
This paper addresses the challenge of quantifying uncertainty in topic–metadata relationships within topic modeling, proposing the first end-to-end fully Bayesian framework—replacing conventional hybrid frequentist-Bayesian approaches. Methodologically, it innovatively integrates Beta regression into the method-of-composition framework of Structural Topic Models (STM) and implements full Bayesian inference via Monte Carlo sampling. The framework jointly estimates latent topic structure and robust statistical associations between topics and metadata (e.g., constituency-level demographic and economic variables), markedly improving the validity and interpretability of posterior uncertainty quantification. Empirical evaluation on German parliamentarians’ tweet data demonstrates the model’s ability to uncover heterogeneous, geographically nuanced associations between issue distributions and regional characteristics. By enabling principled uncertainty propagation and coherent Bayesian inference, the approach provides a more reliable foundation for causal reasoning and policy analysis in the social sciences.
📝 Abstract
The objective of advanced topic modeling is not only to explore latent topical structures, but also to estimate relationships between the discovered topics and theoretically relevant metadata. Methods used to estimate such relationships must take into account that the topical structure is not directly observed, but instead being estimated itself in an unsupervised fashion, usually by common topic models. A frequently used procedure to achieve this is the method of composition , a Monte Carlo sampling technique performing multiple repeated linear regressions of sampled topic proportions on metadata covariates. In this paper, we propose two modifications of this approach: First, we substantially refine the existing implementation of the method of composition from the R package stm by replacing linear regression with the more appropriate Beta regression. Second, we provide a fundamental enhancement of the entire estimation framework by substituting the current blending of frequentist and Bayesian methods with a fully Bayesian approach. This allows for a more appropriate quantification of uncertainty. We illustrate our improved methodology by investigating relationships between Twitter posts by German parliamentarians and different metadata covariates related to their electoral districts, using the structural topic model to estimate topic proportions.
Problem

Research questions and friction points this paper is trying to address.

Estimating relationships between latent topics and metadata
Improving method of composition with Beta regression
Enhancing uncertainty quantification via fully Bayesian approach
Innovation

Methods, ideas, or system contributions that make the work stand out.

Replacing linear regression with Beta regression
Implementing a fully Bayesian approach
Enhancing uncertainty quantification in topic modeling
🔎 Similar Papers
2024-04-02North American Chapter of the Association for Computational LinguisticsCitations: 2
P
P. Schulze
Department of Statistics, Ludwig-Maximilians-Universität, Munich, Germany
Simon Wiegrebe
Simon Wiegrebe
PhD Student in Statistics, LMU Munich
survival analysisdeep learningreduction techniquesmulti-state modelslongitudinal modeling
P
Paul W. Thurner
Department of Political Science, Ludwig-Maximilians-Universität, Munich, Germany
C
C. Heumann
Department of Statistics, Ludwig-Maximilians-Universität, Munich, Germany
M
M. Aßenmacher
Department of Statistics, Ludwig-Maximilians-Universität, Munich, Germany
S
Sandra WankmĂźller