Score
Design and implement probabilistic latent-variable training objectives and encoders that use a variational approximation to the information bottleneck—typically a stochastic encoder plus a KL-divergence term to a prior—to limit mutual information between inputs and learned latent codes while preserving task-relevant signals. Build, apply, and analyze these bottleneck regularizers to produce compact, robust representations and to study how bottleneck strength and priors affect representation compactness, robustness to corruption, and downstream performance, often without changing the base model architecture.
This paper addresses the lack of a unified theoretical framework for variational dimensionality reduction. It proposes a unified Variational Information Bottleneck (VIB) framework that jointly optimizes encoder-based information compression and decoder-based generative fidelity, enabling principled information trade-offs in latent space. Key contributions include: (1) introducing DVSIB and beta-DVCCA—novel methods that extend the multivariate information bottleneck to deep variational settings for the first time; (2) establishing theoretical connections between DSIB and contrastive learning approaches (e.g., Barlow Twins) via mutual information regularization; and (3) proposing symmetric and weighted mutual information regularization to support multi-view representation learning and generative modeling. Evaluated on Noisy MNIST and CIFAR-100, the framework achieves significant improvements in classification accuracy, latent dimension efficiency, and sample efficiency, attaining state-of-the-art or superior performance.
Concept Bottleneck Models (CBMs) predict predefined concepts but fail to satisfy the information bottleneck principle—concept prediction capability does not imply concept-exclusive encoding, undermining interpretability and intervention reliability. Method: We identify this fundamental limitation and propose the Minimum Concept Bottleneck Model (MCBM), which enforces each latent variable to retain only the minimal sufficient information for its associated concept via variational information bottleneck regularization. Contribution/Results: MCBM is the first CBM framework to provide theoretical guarantees for concept interventions, ensuring both Bayesian consistency and architectural flexibility. Empirical evaluation across multiple benchmarks demonstrates significant improvements in concept specificity, intervention robustness, and model interpretability. These results validate the critical role of strict information constraints in building trustworthy, concept-based models.
This work investigates theoretical guarantees for representation learning in self-supervised and semi-supervised settings, aiming to balance information compression with predictive power. Framed through the information bottleneck principle, the problem is cast as a rate–distortion optimization, where optimal representations are obtained via soft clustering on a predictive manifold. The authors propose Sketched Isotropic Gaussian Regularization (SIGReg), which constructs an exact transformation chain from the probability simplex to an isotropic Gaussian distribution, yielding a tractable, non-variational encoder loss. Theoretical analysis combines conditional entropy bottleneck decomposition with minibatch-based marginal estimation. Empirical validation on synthetic data and FashionMNIST demonstrates the effectiveness of the rate–distortion trade-off, with the non-parametric implementation achieving performance comparable to standard variational methods.
This study addresses the limitation of the standard Information Bottleneck (IB) in disentangling label-relevant structures from irrelevant noise, which renders models prone to overfitting in few-shot scenarios. Building upon a label-induced partitioned reconstruction IB, this work proposes a dual-bottleneck framework that independently regulates global capacity and intra-conditional information. By achieving an exact decomposition of the conditional KL divergence and introducing a simplex structural prior to constrain latent space geometry, the method effectively disentangles noise. This approach integrates IB theory, structured latent variable modeling, and deep learning regularization techniques. It yields substantial improvements on low-data classification tasks while maintaining consistent performance gains across dense prediction benchmarks.
The Information Bottleneck (IB) principle suffers from severe overfitting and performance degradation under label noise due to its reliance on exact ground-truth labels. To address this, we propose LaT-IB, a label-noise-robust IB learning framework grounded in the novel “Minimal–Sufficient–Clean” (MSC) principle, which theoretically guarantees separation of clean label information from noise components. LaT-IB introduces a noise-aware latent disentanglement mechanism and a three-stage progressive training strategy—Warmup, Knowledge Injection, and Robust Optimization—integrated with mutual information regularization and disentangled representation learning. Extensive experiments across diverse noise settings demonstrate that LaT-IB consistently outperforms existing IB-based and robust learning methods, achieving significant improvements in classification accuracy and generalization stability. These results validate its effectiveness and practicality in real-world noisy-label scenarios.
This study investigates the Gaussian information bottleneck generalized to jointly stable random variables. Focusing on the additive model X = Y + A, it establishes, for the first time, a theoretical framework for the linear information bottleneck with stable variables. By integrating information theory, probability statistics, and optimization theory, closed-form solutions and critical values for stochastic linear encoders are derived. The analysis demonstrates that such encoders are strictly suboptimal in non-Gaussian settings yet asymptotically optimal under high compression rates. This work not only successfully recovers classical Gaussian information bottleneck results but also reveals optimality boundaries in non-Gaussian scenarios, thereby providing a rigorous theoretical foundation for information compression under stable distributions.
This work addresses the high computational cost and sensitivity to input noise inherent in traditional capsule networks due to iterative dynamic routing. The authors introduce, for the first time, the information bottleneck principle into capsule networks and propose a one-shot variational aggregation mechanism that eliminates iterative routing altogether. By leveraging global context compression and class-specific variational autoencoders, the method directly infers latent capsules in a single pass. This approach substantially enhances both efficiency and robustness: on benchmarks such as MNIST, it achieves an average accuracy improvement of over 14% under noisy conditions, accelerates training by 2.54×, increases inference throughput by 3.64×, reduces parameter count by 4.66%, and maintains high accuracy on clean data.
This work addresses the privacy leakage risk in deep learning inference arising from unauthorized reuse of input data by unintended models for other tasks. It proposes a model-specific representation learning approach that operates without pixel-level reconstruction loss. Built upon a variational autoencoding framework, the method integrates task-driven cross-entropy supervision with KL regularization and introduces a gradient saliency-guided dynamic binary mask to selectively suppress latent dimensions irrelevant to the target classification task. Evaluated on CIFAR-100, the approach maintains high accuracy for the designated classifier while reducing the accuracy of all non-target classifiers to below 2%, achieving a suppression ratio exceeding 45×. The method demonstrates strong generalization across multiple datasets and, to the best of our knowledge, is the first to effectively prevent cross-model transfer of feature representations without relying on reconstruction loss.
This work addresses the challenge in variational autoencoders (VAEs) of simultaneously achieving high representational capacity and disentangled, low-dimensional latent representations. The authors formulate VAE training as a soft-constrained optimization problem, introducing an entropy-based soft constraint mechanism to regulate the information content of individual latent variables. Coupled with weight filtering, this approach enables automatic pruning of low-entropy dimensions. The proposed method enhances representation efficiency while preserving disentanglement. Experiments demonstrate significant improvements: on dSprites, activation scores increase by 43–62%, FactorVAE score reaches 0.891, and reconstruction error decreases by 38%; on MNIST, over 90% classification accuracy is achieved using only two latent dimensions—reducing input dimensionality by 80% compared to baselines—and training convergence accelerates by 37%.
This study addresses the closed-form expression of the Kullback–Leibler (KL) divergence between Gaussian priors and posteriors in variational autoencoders (VAEs) and its role in training dynamics. Drawing on information theory and probability, the work systematically derives analytical solutions for the KL divergence under both univariate and diagonal-covariance multivariate Gaussian assumptions. It further elucidates how individual terms in the KL divergence contribute to regularization in the latent space and shape the model’s generative capacity. By providing a clear and rigorous theoretical derivation, this research deepens the understanding of the intrinsic nature of the VAE’s regularization term and offers principled guidance for model design and practical optimization.