variational information bottleneck

Design and implement probabilistic latent-variable training objectives and encoders that use a variational approximation to the information bottleneck—typically a stochastic encoder plus a KL-divergence term to a prior—to limit mutual information between inputs and learned latent codes while preserving task-relevant signals. Build, apply, and analyze these bottleneck regularizers to produce compact, robust representations and to study how bottleneck strength and priors affect representation compactness, robustness to corruption, and downstream performance, often without changing the base model architecture.

variationalinformationbottleneck

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.2
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This paper addresses the lack of a unified theoretical framework for variational dimensionality reduction. It proposes a unified Variational Information Bottleneck (VIB) framework that jointly optimizes encoder-based information compression and decoder-based generative fidelity, enabling principled information trade-offs in latent space. Key contributions include: (1) introducing DVSIB and beta-DVCCA—novel methods that extend the multivariate information bottleneck to deep variational settings for the first time; (2) establishing theoretical connections between DSIB and contrastive learning approaches (e.g., Barlow Twins) via mutual information regularization; and (3) proposing symmetric and weighted mutual information regularization to support multi-view representation learning and generative modeling. Evaluated on Noisy MNIST and CIFAR-100, the framework achieves significant improvements in classification accuracy, latent dimension efficiency, and sample efficiency, attaining state-of-the-art or superior performance.

Balancing information in encoder and decoder graphsImproving latent spaces for multi-view representation learningUnifying framework for variational dimensionality reduction methods

There Was Never a Bottleneck in Concept Bottleneck Models

Jun 05, 2025
AA
Antonio Almud'evar
🏛️ University of Zaragoza | University of Cambdridge

Concept Bottleneck Models (CBMs) predict predefined concepts but fail to satisfy the information bottleneck principle—concept prediction capability does not imply concept-exclusive encoding, undermining interpretability and intervention reliability. Method: We identify this fundamental limitation and propose the Minimum Concept Bottleneck Model (MCBM), which enforces each latent variable to retain only the minimal sufficient information for its associated concept via variational information bottleneck regularization. Contribution/Results: MCBM is the first CBM framework to provide theoretical guarantees for concept interventions, ensuring both Bayesian consistency and architectural flexibility. Empirical evaluation across multiple benchmarks demonstrates significant improvements in concept specificity, intervention robustness, and model interpretability. These results validate the critical role of strict information constraints in building trustworthy, concept-based models.

CBMs lack true bottleneck, risking interpretability and intervention validityMCBMs enable guaranteed concept interventions with Bayesian consistencyMCBMs use IB to enforce concept-specific information retention

This work investigates theoretical guarantees for representation learning in self-supervised and semi-supervised settings, aiming to balance information compression with predictive power. Framed through the information bottleneck principle, the problem is cast as a rate–distortion optimization, where optimal representations are obtained via soft clustering on a predictive manifold. The authors propose Sketched Isotropic Gaussian Regularization (SIGReg), which constructs an exact transformation chain from the probability simplex to an isotropic Gaussian distribution, yielding a tractable, non-variational encoder loss. Theoretical analysis combines conditional entropy bottleneck decomposition with minibatch-based marginal estimation. Empirical validation on synthetic data and FashionMNIST demonstrates the effectiveness of the rate–distortion trade-off, with the non-parametric implementation achieving performance comparable to standard variational methods.

Distributional RegularizationInformation BottleneckRate-Distortion Trade-off

This study addresses the limitation of the standard Information Bottleneck (IB) in disentangling label-relevant structures from irrelevant noise, which renders models prone to overfitting in few-shot scenarios. Building upon a label-induced partitioned reconstruction IB, this work proposes a dual-bottleneck framework that independently regulates global capacity and intra-conditional information. By achieving an exact decomposition of the conditional KL divergence and introducing a simplex structural prior to constrain latent space geometry, the method effectively disentangles noise. This approach integrates IB theory, structured latent variable modeling, and deep learning regularization techniques. It yields substantial improvements on low-data classification tasks while maintaining consistent performance gains across dense prediction benchmarks.

GeneralizationInformation BottleneckLabel-Induced Partitions

Is the Information Bottleneck Robust Enough? Towards Label-Noise Resistant Information Bottleneck Learning

Dec 11, 2025
YH
Yi Huang
🏛️ Beihang University | HKUST | Guangxi Normal University

The Information Bottleneck (IB) principle suffers from severe overfitting and performance degradation under label noise due to its reliance on exact ground-truth labels. To address this, we propose LaT-IB, a label-noise-robust IB learning framework grounded in the novel “Minimal–Sufficient–Clean” (MSC) principle, which theoretically guarantees separation of clean label information from noise components. LaT-IB introduces a noise-aware latent disentanglement mechanism and a three-stage progressive training strategy—Warmup, Knowledge Injection, and Robust Optimization—integrated with mutual information regularization and disentangled representation learning. Extensive experiments across diverse noise settings demonstrate that LaT-IB consistently outperforms existing IB-based and robust learning methods, achieving significant improvements in classification accuracy and generalization stability. These results validate its effectiveness and practicality in real-world noisy-label scenarios.

Addresses Information Bottleneck's vulnerability to label noiseEnhances robustness in real-world scenarios with noisy labelsProposes a noise-resistant method with latent disentanglement

Latest Papers

What's happening recently
View more

This study investigates the Gaussian information bottleneck generalized to jointly stable random variables. Focusing on the additive model X = Y + A, it establishes, for the first time, a theoretical framework for the linear information bottleneck with stable variables. By integrating information theory, probability statistics, and optimization theory, closed-form solutions and critical values for stochastic linear encoders are derived. The analysis demonstrates that such encoders are strictly suboptimal in non-Gaussian settings yet asymptotically optimal under high compression rates. This work not only successfully recovers classical Gaussian information bottleneck results but also reveals optimality boundaries in non-Gaussian scenarios, thereby providing a rigorous theoretical foundation for information compression under stable distributions.

Gaussian Information BottleneckInformation BottleneckLinear Encoders

This work addresses the high computational cost and sensitivity to input noise inherent in traditional capsule networks due to iterative dynamic routing. The authors introduce, for the first time, the information bottleneck principle into capsule networks and propose a one-shot variational aggregation mechanism that eliminates iterative routing altogether. By leveraging global context compression and class-specific variational autoencoders, the method directly infers latent capsules in a single pass. This approach substantially enhances both efficiency and robustness: on benchmarks such as MNIST, it achieves an average accuracy improvement of over 14% under noisy conditions, accelerates training by 2.54×, increases inference throughput by 3.64×, reduces parameter count by 4.66%, and maintains high accuracy on clean data.

Capsule NetworksComputational CostDynamic Routing

This work addresses the privacy leakage risk in deep learning inference arising from unauthorized reuse of input data by unintended models for other tasks. It proposes a model-specific representation learning approach that operates without pixel-level reconstruction loss. Built upon a variational autoencoding framework, the method integrates task-driven cross-entropy supervision with KL regularization and introduces a gradient saliency-guided dynamic binary mask to selectively suppress latent dimensions irrelevant to the target classification task. Evaluated on CIFAR-100, the approach maintains high accuracy for the designated classifier while reducing the accuracy of all non-target classifiers to below 2%, achieving a suppression ratio exceeding 45×. The method demonstrates strong generalization across multiple datasets and, to the best of our knowledge, is the first to effectively prevent cross-model transfer of feature representations without relying on reconstruction loss.

cross-model transferfeature compressioninput repurposing

This work addresses the challenge in variational autoencoders (VAEs) of simultaneously achieving high representational capacity and disentangled, low-dimensional latent representations. The authors formulate VAE training as a soft-constrained optimization problem, introducing an entropy-based soft constraint mechanism to regulate the information content of individual latent variables. Coupled with weight filtering, this approach enables automatic pruning of low-entropy dimensions. The proposed method enhances representation efficiency while preserving disentanglement. Experiments demonstrate significant improvements: on dSprites, activation scores increase by 43–62%, FactorVAE score reaches 0.891, and reconstruction error decreases by 38%; on MNIST, over 90% classification accuracy is achieved using only two latent dimensions—reducing input dimensionality by 80% compared to baselines—and training convergence accelerates by 37%.

disentanglementencoding capacitylatent space

This study addresses the closed-form expression of the Kullback–Leibler (KL) divergence between Gaussian priors and posteriors in variational autoencoders (VAEs) and its role in training dynamics. Drawing on information theory and probability, the work systematically derives analytical solutions for the KL divergence under both univariate and diagonal-covariance multivariate Gaussian assumptions. It further elucidates how individual terms in the KL divergence contribute to regularization in the latent space and shape the model’s generative capacity. By providing a clear and rigorous theoretical derivation, this research deepens the understanding of the intrinsic nature of the VAE’s regularization term and offers principled guidance for model design and practical optimization.

closed-form expressionGaussian distributionsKL divergence

Hot Scholars

MR

Mickael Rouvier

University of Avignon - LIA
Automatic Speech RecognitionSpeaker DiarizationSpeaker Verification
NE

Nicholas Evans

Professor, Audio Security and Privacy, EURECOM, France
speaker recognitionanti-spoofingpresentation attack detectionprivacy preservation
JW

Jingjing Wang

Professor, School of Cyber Science and Technology, Beihang University
AI for WirelessUAV NetworksSpace-Air-Ground-Sea NetworksCommunication Security
YD

Yiqin Deng

City University of Hong Kong
UAV-enabled Computing Power NetworksResource Scheduling in Edge ComputingEdge AI