Score
Designs and implements generation and sampling procedures that use minimum description length (MDL) scores—e.g., description-length based accept/reject or weighting—to enforce constraints such as novelty, bounded influence on mixture components, or limits on risk. Builds the scoring and thresholding rules, constrained samplers or mixture-aware sampling strategies, and analyzes the resulting probabilistic guarantees (for example bounds on misclassification or on sample impact to model structure) that follow from the MDL-based constraints.
Classical Minimum Description Length (MDL) theory relies on exact optimization, yet practical models can only approximately optimize the objective, leading to a gap between theory and practice. This work presents the first systematic study of predictive reliability under additive approximation errors and regularization in approximate MDL. We introduce a balanced MDL objective and combine additive slack analysis, affinity-telescoping arguments, and a likelihood-ratio-based stopping-time technique. We prove that when the regularization strength λ ≥ 1, the cumulative prediction error remains bounded; conversely, if λ < 1, overfitting is inevitable and multiplicative approximation becomes infeasible. Our results demonstrate that classical MDL is robust to any fixed additive optimization error, and this condition is theoretically tight.
This paper systematically characterizes overfitting behavior along the full regularization path of two-part-code Minimum Description Length (MDL)-based learning rules for binary classification. Adopting the agnostic PAC framework and asymptotic analysis, it derives, for the first time, an explicit limiting expression for the generalization error as a function of the regularization parameter λ and noise level, and establishes tight worst-case upper bounds. The main contributions are threefold: (1) correcting and significantly strengthening GL’s earlier conclusion on non-uniformity at λ = 1; (2) revealing that under-regularization risk exhibits a continuous spectrum—rather than a binary threshold—across λ; and (3) proving that overfitting severity can be continuously and controllably tuned by λ, refuting the conventional “phase-transition” view. Collectively, these results provide the first comprehensive theoretical characterization and quantitative control principle for MDL regularization across the entire λ domain.
Neural networks such as Transformers lack theoretically grounded measures of model complexity, hindering principled model selection and compression. Method: Grounded in the Minimum Description Length (MDL) principle and Kolmogorov complexity, we establish the first asymptotically optimal description length objective for Transformers and construct the first MDL framework with computational universality guarantees. We propose a differentiable, optimization-friendly variational objective using an adaptive Gaussian mixture prior to approximate MDL. Contribution/Results: This work introduces the first theoretically sound, Transformer-specific MDL-based complexity measure. Empirical evaluation confirms that the proposed objective favors low-complexity models with strong generalization performance. However, it also exposes a critical practical limitation: standard optimizers struggle to converge from random initialization. Overall, our framework provides a novel information-theoretic foundation for model selection and compression in deep learning, bridging theoretical guarantees with practical neural architecture design.
In Bayesian sequential trials, error rate evaluation relies on computationally expensive Monte Carlo simulations, hindering efficient optimization of sample size and decision thresholds. Method: This paper establishes, for the first time, analytical functional relationships between posterior and posterior predictive probabilities and sample size. Leveraging Bayesian decision theory and asymptotic analysis—combined with numerical fitting and error-rate inversion—the method enables precise error-rate assessment for any sample size using only two simulations, and rapidly identifies optimal design parameters. Contribution/Results: The approach drastically reduces computational cost while achieving error-rate control accuracy comparable to conventional simulation-based methods. In two real-world case studies, it attains exact error-rate calibration and accelerates design optimization by several orders of magnitude. This provides a scalable, verifiable, and highly efficient design paradigm for Bayesian adaptive trials.
Subspace clustering in high-dimensional data often yields multiple semantically distinct subspaces, yet existing methods require manual specification of both the number of subspaces and the number of clusters within each—rendering them parameter-sensitive and poorly interpretable. This paper proposes an automatic, non-redundant multi-subspace clustering framework. First, it introduces the Minimum Description Length (MDL) principle to non-redundant clustering, enabling joint, adaptive inference of both the optimal number of subspaces and the cluster count per subspace. Second, it designs a split-merge-based greedy search strategy coupled with a subspace-level outlier encoding mechanism, allowing simultaneous outlier detection. Evaluated on multiple benchmark datasets, the method achieves competitive accuracy against state-of-the-art approaches while significantly improving parameter robustness, model interpretability, and practical applicability.
本文探讨了贝叶斯样本量确定的两种方法:估计抽样分布或通过随机根查找探索,评估了它们在复杂模型中的性能。
This work addresses the challenge of risk prediction when acquiring true outcomes is prohibitively expensive, allowing labels for only a subset of samples. The authors propose a surrogate-assisted optimal sampling framework that, under a fixed annotation budget, leverages covariates, surrogate variables, and an initial estimator to construct a sampling strategy minimizing the expected out-of-sample cross-entropy loss. Coupled with an inverse probability weighted cross-entropy estimator for model training, this approach achieves—without requiring access to true responses during design—theoretically optimal sampling. It uniquely guarantees predictive optimality, robustness to surrogate misspecification, and stability in settings with rare outcomes. Both theoretical analysis and empirical experiments demonstrate that the method significantly outperforms existing approaches, particularly when the surrogate is imperfect or the event of interest is rare.
This work addresses the inefficiency of existing sampling-based inference methods, which struggle to effectively explore critical decision points during resampling. The authors propose a training-free inference optimization approach that identifies significant entropy spikes in the base model’s next-token prediction along the reasoning trajectory as proxies for key decision points. By integrating an entropy-guided cutoff selection strategy with the Metropolis-Hastings sampling algorithm, the method enables targeted resampling that substantially reduces mixing time complexity. Evaluated on multiple challenging benchmarks—including MATH500, HumanEval, GPQA Diamond, and AIME26—the approach consistently outperforms current baselines and even reinforcement learning–trained models, demonstrating strong training-free reasoning capabilities.
While test-time scaling can enhance the reasoning performance of large language models, it incurs substantial computational overhead and latency. Existing adaptive sampling methods often rely on heuristic rules or distributional assumptions, making it challenging to efficiently balance accuracy and cost. This work formulates adaptive sampling as a Markov decision process and introduces the first lightweight reinforcement learning–based controller that operates solely on final-answer statistics, enabling training and deployment entirely on CPU. By incorporating Lagrangian relaxation, the approach achieves joint optimization under explicit budget constraints. Experiments demonstrate that the proposed method significantly improves the trade-off among accuracy, number of sampling rounds, and total sample count across multiple strong baselines, including ASC and ESC.
本文探讨了在场景优化和无分布认证中,通过确定性边界机制及随机观察边界大小来精确计算违反风险的方法。