Last Layer Hamiltonian Monte Carlo

📅 2025-07-11
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
To address the challenge of efficient uncertainty quantification for large-scale deep neural networks under resource-constrained settings, this paper proposes Last-layer Hamiltonian Monte Carlo (LL-HMC): a method that applies HMC sampling exclusively to the final-layer weights, drastically reducing computational overhead compared to full-parameter HMC. The approach supports multi-chain parallel sampling and grid-search-based hyperparameter optimization. Evaluated on three real-world video datasets for driver behavior and intention recognition, LL-HMC achieves classification accuracy and calibration performance comparable to state-of-the-art last-layer probabilistic methods (e.g., LL-BNN, LL-SVGD), while significantly outperforming them in out-of-distribution (OOD) detection. Empirical analysis further reveals that increasing the number of sampling chains enhances OOD discrimination capability but yields no improvement in classification accuracy. This work establishes a new paradigm for high-confidence, low-overhead uncertainty modeling at the edge.

Technology Category

Machine Learning: Calibration & Uncertainty QuantificationReasoning under Uncertainty: Stochastic OptimizationSearch and Optimization: Sampling/Simulation-based Search

Application Category

Economics, Online Markets and Human Computation: Cost models of using LLMs in production systemsGraph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphsResponsible Web: Machine-in-the-loop, human agency and autonomy
📝 Abstract
We explore the use of Hamiltonian Monte Carlo (HMC) sampling as a probabilistic last layer approach for deep neural networks (DNNs). While HMC is widely regarded as a gold standard for uncertainty estimation, the computational demands limit its application to large-scale datasets and large DNN architectures. Although the predictions from the sampled DNN parameters can be parallelized, the computational cost still scales linearly with the number of samples (similar to an ensemble). Last layer HMC (LL--HMC) reduces the required computations by restricting the HMC sampling to the final layer of a DNN, making it applicable to more data-intensive scenarios with limited computational resources. In this paper, we compare LL-HMC against five last layer probabilistic deep learning (LL-PDL) methods across three real-world video datasets for driver action and intention. We evaluate the in-distribution classification performance, calibration, and out-of-distribution (OOD) detection. Due to the stochastic nature of the probabilistic evaluations, we performed five grid searches for different random seeds to avoid being reliant on a single initialization for the hyperparameter configurations. The results show that LL--HMC achieves competitive in-distribution classification and OOD detection performance. Additional sampled last layer parameters do not improve the classification performance, but can improve the OOD detection. Multiple chains or starting positions did not yield consistent improvements.
Problem

Research questions and friction points this paper is trying to address.

Reducing computational cost of HMC for large DNNs
Applying HMC only to last layer for scalability
Evaluating LL-HMC vs other methods on real-world datasets
Innovation

Methods, ideas, or system contributions that make the work stand out.

HMC sampling in last DNN layer
Reduces computational cost significantly
Competitive performance in classification
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.