Score
Design discrete bottlenecks: define and implement architectural components that map continuous inputs to discrete latent representations (e.g., categorical latents, vector‑quantized codebooks) with specified quantization schemes, codebook sizes, and capacity constraints, and construct the training objectives and routing mechanisms that operate around them. Analyze and optimize how these bottlenecks control information flow—measuring and reducing quantization error, preventing unwanted entanglement of factors, and enabling conditional isolation or routing of residual signals to achieve the desired tradeoff between compression and fidelity.
Deploying deep neural networks (DNNs) faces challenges from high computational overhead and large model sizes. While low-bit weight quantization accelerates inference and reduces memory bandwidth requirements, it often incurs substantial accuracy degradation. This paper presents a systematic survey of low-bit weight quantization research from 2019 to 2024. We propose the first unified taxonomy comprising eight major categories and 24 subcategories—covering linear/nonlinear quantization, layer-wise/channel-wise calibration, retraining-free and fine-tuning-based paradigms, gradient approximation techniques, and mixed-precision search strategies. Through structured comparative analysis of over 100 state-of-the-art works, we identify common bottlenecks, clarify promising future directions, and highlight open challenges. To foster reproducibility and industrial adoption, we open-source Awesome-Model-Quantization—a curated, continuously updated resource repository—thereby advancing standardization and practical deployment of quantization techniques.
Concept Bottleneck Models (CBMs) predict predefined concepts but fail to satisfy the information bottleneck principle—concept prediction capability does not imply concept-exclusive encoding, undermining interpretability and intervention reliability. Method: We identify this fundamental limitation and propose the Minimum Concept Bottleneck Model (MCBM), which enforces each latent variable to retain only the minimal sufficient information for its associated concept via variational information bottleneck regularization. Contribution/Results: MCBM is the first CBM framework to provide theoretical guarantees for concept interventions, ensuring both Bayesian consistency and architectural flexibility. Empirical evaluation across multiple benchmarks demonstrates significant improvements in concept specificity, intervention robustness, and model interpretability. These results validate the critical role of strict information constraints in building trustworthy, concept-based models.
This paper investigates the fundamental differences in generative diversity among discrete latent generative models—autoregressive (AR), masked image modeling (MIM), and diffusion models. We propose the first diagnostic framework grounded in information bottleneck theory, decomposing diversity into *path diversity* (stochasticity in sampling trajectories) and *execution diversity* (output variability conditioned on a fixed trajectory), and design three zero-shot inference-time intervention methods for empirical analysis. Our findings reveal distinct trade-off strategies: MIM prioritizes diversity, AR favors compression, and diffusion enables decoupled control over path and execution diversity. Consequently, we uncover the underlying compression–diversity trade-off mechanism and introduce a plug-and-play inference-time diversity enhancement technique that significantly improves generative diversity without compromising fidelity.
This work addresses the unreliability of concept–prediction associations in Concept Bottleneck Models (CBMs), often caused by concept leakage and accompanied by degraded accuracy. The authors propose a theoretically grounded, architecture-agnostic information bottleneck regularization method that learns minimal sufficient concept representations by minimizing the mutual information \(I(X;C)\) between inputs and concepts while preserving the mutual information \(I(C;Y)\) between concepts and labels—all without modifying model architecture or requiring additional supervision. By integrating a variational objective with entropy-based proxy constraints, the approach seamlessly fits into standard CBM training pipelines. Information plane analysis confirms its mechanistic efficacy. Evaluated across six CBM variants and three benchmark datasets, the method consistently improves prediction accuracy, mitigates concept leakage, and enhances the stability of concept interventions.
Existing concept bottleneck models (CBMs) struggle to capture high-order interactions among concepts and cannot quantify the conditional dependence probabilities between concepts and predictions, limiting their interpretability and causal intervention capability. To address these limitations, we propose the Energy-driven Concept Bottleneck Model (ECBM), the first CBM that formulates the concept bottleneck as a differentiable joint energy function over inputs, concepts, and labels. This unified representation explicitly encodes nonlinear concept interactions and cross-level conditional dependencies, thereby overcoming the restrictive independence assumptions and intervention failures inherent in conventional CBMs. Our method integrates neural energy function design, energy decomposition–based conditional probability derivation, contrastive learning, and MCMC-based approximate inference. Evaluated on multiple real-world datasets, ECBM achieves significant improvements over state-of-the-art methods in classification accuracy, while enabling computable concept correction paths and fine-grained probabilistic explanations.
This paper addresses the lack of a unified theoretical framework for variational dimensionality reduction. It proposes a unified Variational Information Bottleneck (VIB) framework that jointly optimizes encoder-based information compression and decoder-based generative fidelity, enabling principled information trade-offs in latent space. Key contributions include: (1) introducing DVSIB and beta-DVCCA—novel methods that extend the multivariate information bottleneck to deep variational settings for the first time; (2) establishing theoretical connections between DSIB and contrastive learning approaches (e.g., Barlow Twins) via mutual information regularization; and (3) proposing symmetric and weighted mutual information regularization to support multi-view representation learning and generative modeling. Evaluated on Noisy MNIST and CIFAR-100, the framework achieves significant improvements in classification accuracy, latent dimension efficiency, and sample efficiency, attaining state-of-the-art or superior performance.
This study addresses the limitation of the standard Information Bottleneck (IB) in disentangling label-relevant structures from irrelevant noise, which renders models prone to overfitting in few-shot scenarios. Building upon a label-induced partitioned reconstruction IB, this work proposes a dual-bottleneck framework that independently regulates global capacity and intra-conditional information. By achieving an exact decomposition of the conditional KL divergence and introducing a simplex structural prior to constrain latent space geometry, the method effectively disentangles noise. This approach integrates IB theory, structured latent variable modeling, and deep learning regularization techniques. It yields substantial improvements on low-data classification tasks while maintaining consistent performance gains across dense prediction benchmarks.
Directly applying language gradients to discrete symbolic representations in world models often leads to symbol collapse and failed semantic binding. This work proposes a write-protected discrete bottleneck architecture that safely interfaces language with world models without fine-tuning the underlying large language model and with fewer than 2 million additional parameters. The approach combines gradient blocking (via z.detach()), a gradient-free semantic channel implemented as a non-parametric memory table based on symbol–label co-occurrence counts, and DP-Means streaming clustering to dynamically resolve conflicts. It is the first method to simultaneously preserve symbol diversity and achieve high-fidelity semantic binding, attaining zero collapse across 74 experiments and binding accuracies of 79–100% (peaking at 97.2%), substantially outperforming baselines by 22.2%, while remaining compatible with diverse encoders and environments.
This study addresses the long-standing limitation in the Information Bottleneck (IB) framework, where the cardinality bound for optimal representations of binary sources has been constrained by generic upper bounds, thereby hindering computational efficiency. By exploiting the structural properties specific to the binary case, this work employs a separating hyperplane argument combined with concavity analysis of the ratio of second derivatives of entropy functions to transcend traditional generic bounds. It rigorously proves that the optimal representation for a binary source is itself binary, tightening the classical cardinality bound to the exact limit |U|≤|X|. This contribution not only establishes a theoretically optimal bound but also substantially reduces the computational complexity of solving IB problems.
This study addresses the challenge of preserving predictive information about a target variable while removing irrelevant redundancy in data compression. Building on statistical decision theory, the authors propose an ℋ-mutual information framework that satisfies conditional independence (CV) and average generalization (AVG) criteria. They establish, for the first time, an equivalence between the generalized information bottleneck problem and Expected Sample Information (ESI), thereby enabling a computable characterization of a representation’s predictive utility. An alternating optimization algorithm is further developed to efficiently approximate the Pareto frontier between compression and utility. This work extends the applicability of classical mutual information and offers a new paradigm for information bottleneck theory that balances theoretical rigor with practical utility.
This work addresses the limitations of traditional Concept Bottleneck Models (CBMs), which rely on manually predefined concepts that often lack task relevance or learnability, thereby constraining model performance. To overcome this, the authors propose the Mechanistic Concept Bottleneck Model (M-CBM), which automatically extracts task-relevant, interpretable latent concepts directly from the internal representations of a black-box model using sparse autoencoders. These extracted concepts are then semantically labeled and annotated via a multimodal large language model to construct an adaptive bottleneck layer. For fair evaluation of information leakage and interpretability, the study introduces Normalized Concept Consistency (NCC), a decision-level sparsity metric. Experiments across multiple datasets demonstrate that M-CBM significantly outperforms existing CBM approaches under matched sparsity levels, achieving higher concept prediction accuracy while providing concise and human-interpretable decision rationales.