Score
Designs, builds, or analyzes representation-learning objectives and regularizers based on the information bottleneck principle that constrain the mutual information between inputs and learned representations to compress observations while preserving information relevant to a target variable. In the complementary variant, designs mechanisms that split or coordinate multiple representations or channels so that each retains complementary task-relevant information and suppresses task‑irrelevant input features (i.e., implements and studies IB regularization and complementary IB objectives and their trade‑offs).
This study addresses the challenge of preserving predictive information about a target variable while removing irrelevant redundancy in data compression. Building on statistical decision theory, the authors propose an ℋ-mutual information framework that satisfies conditional independence (CV) and average generalization (AVG) criteria. They establish, for the first time, an equivalence between the generalized information bottleneck problem and Expected Sample Information (ESI), thereby enabling a computable characterization of a representation’s predictive utility. An alternating optimization algorithm is further developed to efficiently approximate the Pareto frontier between compression and utility. This work extends the applicability of classical mutual information and offers a new paradigm for information bottleneck theory that balances theoretical rigor with practical utility.
Concept Bottleneck Models (CBMs) predict predefined concepts but fail to satisfy the information bottleneck principle—concept prediction capability does not imply concept-exclusive encoding, undermining interpretability and intervention reliability. Method: We identify this fundamental limitation and propose the Minimum Concept Bottleneck Model (MCBM), which enforces each latent variable to retain only the minimal sufficient information for its associated concept via variational information bottleneck regularization. Contribution/Results: MCBM is the first CBM framework to provide theoretical guarantees for concept interventions, ensuring both Bayesian consistency and architectural flexibility. Empirical evaluation across multiple benchmarks demonstrates significant improvements in concept specificity, intervention robustness, and model interpretability. These results validate the critical role of strict information constraints in building trustworthy, concept-based models.
The Information Bottleneck (IB) theory offers valuable insights into neural network learning but suffers from theoretical ambiguity and practical intractability due to the difficulty of estimating mutual information. To address these limitations, we propose the Generalized Information Bottleneck (GIB) framework, which replaces mutual information with interaction information (II) and introduces average interaction information as a computationally tractable, cooperative measure of representation synergy. GIB reformulates the IB objective—balancing compression and prediction—while preserving theoretical compatibility with classical IB. Crucially, GIB significantly improves estimability and broad applicability across diverse architectures. Empirical evaluation demonstrates that GIB consistently captures sharp compression phase transitions during training in ReLU networks, CNNs, and Transformers. Moreover, the learned representations exhibit strong alignment with model adversarial robustness. Overall, GIB provides a more rigorous, interpretable, and practically deployable theoretical foundation for information-theoretic analysis of deep learning.
Traditional information bottleneck (IB) methods suffer from insufficient feature representation and optimization drift due to reliance on a fragile variational lower bound and a single encoder. To address this, we propose a Structured IB framework that introduces an auxiliary encoder to explicitly model discriminative, structured features overlooked by the primary encoder—enabling complementary latent-space representations and task-aware information distillation. This work is the first to embed structured feature learning into the IB paradigm, eliminating dependence on strong architectural assumptions. Evaluated across multiple benchmark tasks, our method achieves significant improvements in prediction accuracy, reduces model parameters by 23%, and increases task-relevant mutual information retention by 31%. These results demonstrate a synergistic enhancement of information completeness and generalization capability.
Deep neural networks lack biologically plausible selective attention mechanisms, limiting both efficiency and accuracy in image recognition. To address this, we propose a spatial attention module grounded in information bottleneck theory. Our method explicitly optimizes mutual information: it minimizes the mutual information between the attention representation and the input to suppress redundancy, while maximizing the mutual information between the attention representation and task labels to enhance discriminability. Crucially, we introduce learnable anchors to quantize continuous attention scores—a novel design that strengthens information constraints and improves interpretability of attention maps. By integrating variational attention modeling with deep network embedding, our approach achieves significant performance gains across image classification, fine-grained recognition, and cross-domain classification tasks. The resulting attention maps exhibit high discriminability, strong background suppression, and enhanced interpretability.
This study addresses the long-standing limitation in the Information Bottleneck (IB) framework, where the cardinality bound for optimal representations of binary sources has been constrained by generic upper bounds, thereby hindering computational efficiency. By exploiting the structural properties specific to the binary case, this work employs a separating hyperplane argument combined with concavity analysis of the ratio of second derivatives of entropy functions to transcend traditional generic bounds. It rigorously proves that the optimal representation for a binary source is itself binary, tightening the classical cardinality bound to the exact limit |U|≤|X|. This contribution not only establishes a theoretically optimal bound but also substantially reduces the computational complexity of solving IB problems.
This study addresses the limitation of the standard Information Bottleneck (IB) in disentangling label-relevant structures from irrelevant noise, which renders models prone to overfitting in few-shot scenarios. Building upon a label-induced partitioned reconstruction IB, this work proposes a dual-bottleneck framework that independently regulates global capacity and intra-conditional information. By achieving an exact decomposition of the conditional KL divergence and introducing a simplex structural prior to constrain latent space geometry, the method effectively disentangles noise. This approach integrates IB theory, structured latent variable modeling, and deep learning regularization techniques. It yields substantial improvements on low-data classification tasks while maintaining consistent performance gains across dense prediction benchmarks.
This study addresses the limitation of the classical information bottleneck in directly characterizing downstream decision errors by investigating the Chernoff bottleneck under mutual information constraints, aiming to maximize the error exponent of binary hypothesis testing under rate-limited conditions. Theoretically, it reveals the non-concavity of the Chernoff bottleneck curve and proves that optimality is achievable with only k+1 outputs. Algorithmically, an alternating optimization framework guaranteeing convergence and feasibility is proposed, integrating a generalized Blahut-Arimoto algorithm with nonlinear iterations for joint solution. Experiments on the 20 Newsgroups dataset demonstrate that compressing information while retaining merely 17% of its entropy preserves 90% of the error exponent and achieves near-lossless classification accuracy.
This work investigates theoretical guarantees for representation learning in self-supervised and semi-supervised settings, aiming to balance information compression with predictive power. Framed through the information bottleneck principle, the problem is cast as a rate–distortion optimization, where optimal representations are obtained via soft clustering on a predictive manifold. The authors propose Sketched Isotropic Gaussian Regularization (SIGReg), which constructs an exact transformation chain from the probability simplex to an isotropic Gaussian distribution, yielding a tractable, non-variational encoder loss. Theoretical analysis combines conditional entropy bottleneck decomposition with minibatch-based marginal estimation. Empirical validation on synthetic data and FashionMNIST demonstrates the effectiveness of the rate–distortion trade-off, with the non-parametric implementation achieving performance comparable to standard variational methods.
This work addresses the unreliability of concept–prediction associations in Concept Bottleneck Models (CBMs), often caused by concept leakage and accompanied by degraded accuracy. The authors propose a theoretically grounded, architecture-agnostic information bottleneck regularization method that learns minimal sufficient concept representations by minimizing the mutual information \(I(X;C)\) between inputs and concepts while preserving the mutual information \(I(C;Y)\) between concepts and labels—all without modifying model architecture or requiring additional supervision. By integrating a variational objective with entropy-based proxy constraints, the approach seamlessly fits into standard CBM training pipelines. Information plane analysis confirms its mechanistic efficacy. Evaluated across six CBM variants and three benchmark datasets, the method consistently improves prediction accuracy, mitigates concept leakage, and enhances the stability of concept interventions.