Posterior Approximation using Stochastic Gradient Ascent with Adaptive Stepsize

📅 2024-12-12
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This paper addresses the computational intractability of posterior inference in Dirichlet Process Mixture Models (DPMMs) under large-scale data regimes. We propose a scalable variational inference method based on stochastic gradient ascent. Unlike conventional coordinate ascent relying on closed-form updates, our approach is the first to incorporate an adaptive step-size scheme—integrating momentum and the Fisher information matrix—into stochastic optimization for DPMMs, thereby eliminating dependence on analytical gradient updates. Features are extracted via deep convolutional neural networks, and extensive experiments on large-scale benchmarks—including Caltech256 and SUN397—demonstrate that the proposed method achieves comparable accuracy to coordinate ascent while significantly accelerating training. This enables efficient and robust posterior approximation for Bayesian nonparametric models in large-scale settings.

Technology Category

Machine Learning: Probabilistic Circuits and Graphical ModelsReasoning under Uncertainty: Stochastic OptimizationNatural Language Processing: Learning & Optimization for NLP

Application Category

Graph Algorithms and Modeling for the Web: Graph neural networks and deep learning approaches for Web-related graphsWeb Mining and Content Analysis: Large pretrained models with web dataUser Modeling, Personalization and Recommendation: Practical large-scale studies of user experience
📝 Abstract
Scalable algorithms of posterior approximation allow Bayesian nonparametrics such as Dirichlet process mixture to scale up to larger dataset at fractional cost. Recent algorithms, notably the stochastic variational inference performs local learning from minibatch. The main problem with stochastic variational inference is that it relies on closed form solution. Stochastic gradient ascent is a modern approach to machine learning and is widely deployed in the training of deep neural networks. In this work, we explore using stochastic gradient ascent as a fast algorithm for the posterior approximation of Dirichlet process mixture. However, stochastic gradient ascent alone is not optimal for learning. In order to achieve both speed and performance, we turn our focus to stepsize optimization in stochastic gradient ascent. As as intermediate approach, we first optimize stepsize using the momentum method. Finally, we introduce Fisher information to allow adaptive stepsize in our posterior approximation. In the experiments, we justify that our approach using stochastic gradient ascent do not sacrifice performance for speed when compared to closed form coordinate ascent learning on these datasets. Lastly, our approach is also compatible with deep ConvNet features as well as scalable to large class datasets such as Caltech256 and SUN397.
Problem

Research questions and friction points this paper is trying to address.

Posterior approximation scalability
Stochastic gradient ascent optimization
Adaptive stepsize for performance
Innovation

Methods, ideas, or system contributions that make the work stand out.

Stochastic gradient ascent optimization
Adaptive stepsize using Fisher information
Compatible with deep ConvNet features
🔎 Similar Papers
💼 Related Jobs
No related jobs found.