Score
Designs and implements computable estimators and practical proxies for description length and Kolmogorov complexity (including minimum-description-length criteria) and computes explicit description-length estimates for data, models, or representations. Uses those estimates to compare representation costs and to guide model selection, compression decisions, or complexity-aware analysis.
Neural networks such as Transformers lack theoretically grounded measures of model complexity, hindering principled model selection and compression. Method: Grounded in the Minimum Description Length (MDL) principle and Kolmogorov complexity, we establish the first asymptotically optimal description length objective for Transformers and construct the first MDL framework with computational universality guarantees. We propose a differentiable, optimization-friendly variational objective using an adaptive Gaussian mixture prior to approximate MDL. Contribution/Results: This work introduces the first theoretically sound, Transformer-specific MDL-based complexity measure. Empirical evaluation confirms that the proposed objective favors low-complexity models with strong generalization performance. However, it also exposes a critical practical limitation: standard optimizers struggle to converge from random initialization. Overall, our framework provides a novel information-theoretic foundation for model selection and compression in deep learning, bridging theoretical guarantees with practical neural architecture design.
Classical Minimum Description Length (MDL) theory relies on exact optimization, yet practical models can only approximately optimize the objective, leading to a gap between theory and practice. This work presents the first systematic study of predictive reliability under additive approximation errors and regularization in approximate MDL. We introduce a balanced MDL objective and combine additive slack analysis, affinity-telescoping arguments, and a likelihood-ratio-based stopping-time technique. We prove that when the regularization strength λ ≥ 1, the cumulative prediction error remains bounded; conversely, if λ < 1, overfitting is inevitable and multiplicative approximation becomes infeasible. Our results demonstrate that classical MDL is robust to any fixed additive optimization error, and this condition is theoretically tight.
This study addresses the theoretical assessment of neural network compressibility limits. We extend the Minimum Description Length (MDL) principle—traditionally applicable only to regular models—to singular statistical models by integrating singular learning theory, and propose a novel model complexity estimator based on the Local Learning Coefficient (LLC). Systematic evaluation on the Pythia model family demonstrates a strong linear correlation between LLC and achievable compression ratios across diverse techniques, including quantization and tensor decomposition. Our approach yields the first computationally tractable, theoretically grounded complexity measure for neural networks, overcoming the fundamental limitation that conventional MDL is inapplicable to non-regular (singular) models. By establishing a principled, interpretable link between intrinsic model complexity and compressibility, this work provides a rigorous theoretical framework for characterizing fundamental compression limits in deep learning.
This work investigates the generalization ability and sample complexity of supervised learning from an information-theoretic perspective, modeling the learning process as lossy compression with finite blocklength: training data sampling corresponds to encoding, and model construction to decoding. It introduces, for the first time, finite-blocklength lossy compression theory to analyze generalization error, deriving fundamental lower bounds on both generalization error and sample complexity for any fixed randomized learning algorithm and its optimal sampling strategy. The proposed framework cleanly disentangles the distinct contributions of overfitting and task-inductive bias mismatch, while unifying information-theoretic generalization bounds with the algorithmic stability perspective, thereby revealing their essential roles in determining generalization performance.
Segmented linear approximation (PLA) in learned indexes suffers from suboptimal storage efficiency, and no information-theoretic space lower bound exists for PLA under both compression and indexing constraints. Method: We establish the first information-theoretic space lower bound for PLA in these dual settings, then design a novel, minimalist data structure that achieves theoretically optimal compact representation for 2D monotonic point sequences under a given error bound. The structure supports O(log n)-time x-value lookup and segment evaluation. Contribution/Results: Our approach unifies the modeling of PLA’s compressibility and queryability—yielding the first systematic lower-bound analysis, constructive guarantee, and efficient implementation for PLA-based learned indexes. The space usage is asymptotically tight to the lower bound, achieving succinctness on most practical distributions. This work bridges a critical theoretical and engineering gap in learned indexing research.
This work proposes the concept of *prompt complexity* to quantify the minimal length of a reasonable prompt required to elicit a target text or behavior from a fixed instruction-tuned language model. Inspired by Kolmogorov complexity yet tailored to model-specific interfaces, it formally defines a non-universal, resource-bounded notion of prompt complexity and introduces derived metrics such as prompt distance and behavior-level complexity. The authors develop a computable theoretical framework by integrating readable prompt enumeration under deterministic decoding, soft prompt optimization, text compression theory, and formal specification verification. This framework not only provides a precise optimization objective for prompt engineering but also establishes an empirical research agenda to systematically investigate the accessibility boundaries of desired outputs and behaviors across different models.
本文提出了一种描述复杂性信息准则(DCIC)来解决强预测变量依赖性和模型类别不确定性下的模型选择问题,通过Kraft可接纳码长进行正则化。
论文探讨了如何通过压缩算法转换成机器学习方法,利用归一化压缩距离或最小描述长度原则解决AI中的问题,并提出了一种基于压缩的机器学习设计框架。
该研究通过引入C-RASP+和C-RASP1片段,并采用压缩字符串方法,解决了变换器长度泛化的理论边界问题,提供了一个更紧致的样本大小界限。
This work addresses the limitations of traditional Block Decomposition Method (BDM), which neglects inter-block dependencies, leading to redundant description lengths and an inability to capture shared structures. The authors propose an improved framework grounded in reusable program code, modeling inter-block dependencies as a descriptive resource optimization problem formalized through an “algorithmic attention” allocation mechanism. By integrating Coding Theorem Method (CTM), conditional algorithmic complexity, Shannon entropy, and combinatorial optimization, they construct a computable algorithmic attention model. Theoretically, the approach is shown to strictly outperform independent block descriptions in the presence of shared structures, with performance gains positively correlated to algorithmic mutual information, and the paper provides a feasible implementation pathway for practical application.