🤖 AI Summary
This study addresses the limitation of the classical information bottleneck in directly characterizing downstream decision errors by investigating the Chernoff bottleneck under mutual information constraints, aiming to maximize the error exponent of binary hypothesis testing under rate-limited conditions. Theoretically, it reveals the non-concavity of the Chernoff bottleneck curve and proves that optimality is achievable with only k+1 outputs. Algorithmically, an alternating optimization framework guaranteeing convergence and feasibility is proposed, integrating a generalized Blahut-Arimoto algorithm with nonlinear iterations for joint solution. Experiments on the 20 Newsgroups dataset demonstrate that compressing information while retaining merely 17% of its entropy preserves 90% of the error exponent and achieves near-lossless classification accuracy.
📝 Abstract
The classical information bottleneck (IB) measures the relevance of a representation $U$ of $X$ to a target $Y$ by $I(U;Y)$, which does not directly characterize the error of downstream decisions. For a binary hypothesis $Y$ inferred from many separately encoded observations, the optimal error exponent is the Chernoff information between the two conditional distributions of $U$ given $Y$. We study the mutual information constrained Chernoff bottleneck, which seeks an encoder that maximizes this Chernoff information subject to a rate constraint $I(U;X) \leq R$. We show that its optimal value $C(R)$ increases strictly up to $R = H(V)$, where $V$ merges the symbols of $X$ with equal likelihood ratio, remains at the uncompressed exponent beyond, and, unlike the IB curve, need not be concave. We further show that $k+1$ outputs suffice to attain $C(R)$, where $k$ is the cardinality of $V$. We propose an alternating algorithm that updates the encoder via a generalized Blahut--Arimoto algorithm and the Chernoff parameter $s$ via a nonlinear equation, and prove that its iterates remain feasible, with nondecreasing and convergent Chernoff information. Numerical experiments confirm the theory, and on real topic-detection data from the 20 Newsgroups corpus, compressing each word to only $17\%$ of its entropy retains $90\%$ of the error exponent and nearly the accuracy of the uncompressed classifier.