mdl-constrained generation

Designs and implements generation and sampling procedures that use minimum description length (MDL) scores—e.g., description-length based accept/reject or weighting—to enforce constraints such as novelty, bounded influence on mixture components, or limits on risk. Builds the scoring and thresholding rules, constrained samplers or mixture-aware sampling strategies, and analyzes the resulting probabilistic guarantees (for example bounds on misclassification or on sample impact to model structure) that follow from the MDL-based constraints.

mdl-constrainedgeneration

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.26
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Classical Minimum Description Length (MDL) theory relies on exact optimization, yet practical models can only approximately optimize the objective, leading to a gap between theory and practice. This work presents the first systematic study of predictive reliability under additive approximation errors and regularization in approximate MDL. We introduce a balanced MDL objective and combine additive slack analysis, affinity-telescoping arguments, and a likelihood-ratio-based stopping-time technique. We prove that when the regularization strength λ ≥ 1, the cumulative prediction error remains bounded; conversely, if λ < 1, overfitting is inevitable and multiplicative approximation becomes infeasible. Our results demonstrate that classical MDL is robust to any fixed additive optimization error, and this condition is theoretically tight.

Approximate MDLCompression GuaranteeModel Selection

Quantifying Overfitting along the Regularization Path for Two-Part-Code MDL in Supervised Classification

Mar 03, 2025
XZ
Xiaohan Zhu
🏛️ The University of Chicago | Toyota Technological Institute at Chicago

This paper systematically characterizes overfitting behavior along the full regularization path of two-part-code Minimum Description Length (MDL)-based learning rules for binary classification. Adopting the agnostic PAC framework and asymptotic analysis, it derives, for the first time, an explicit limiting expression for the generalization error as a function of the regularization parameter λ and noise level, and establishes tight worst-case upper bounds. The main contributions are threefold: (1) correcting and significantly strengthening GL’s earlier conclusion on non-uniformity at λ = 1; (2) revealing that under-regularization risk exhibits a continuous spectrum—rather than a binary threshold—across λ; and (3) proving that overfitting severity can be continuously and controllably tuned by λ, refuting the conventional “phase-transition” view. Collectively, these results provide the first comprehensive theoretical characterization and quantitative control principle for MDL regularization across the entire λ domain.

Characterizes regularization curve for MDL in binary classification.Extends analysis of under-regularization and overfitting for all λ values.Quantifies worst-case error based on regularization and noise levels.

Neural networks such as Transformers lack theoretically grounded measures of model complexity, hindering principled model selection and compression. Method: Grounded in the Minimum Description Length (MDL) principle and Kolmogorov complexity, we establish the first asymptotically optimal description length objective for Transformers and construct the first MDL framework with computational universality guarantees. We propose a differentiable, optimization-friendly variational objective using an adaptive Gaussian mixture prior to approximate MDL. Contribution/Results: This work introduces the first theoretically sound, Transformer-specific MDL-based complexity measure. Empirical evaluation confirms that the proposed objective favors low-complexity models with strong generalization performance. However, it also exposes a critical practical limitation: standard optimizers struggle to converge from random initialization. Overall, our framework provides a novel information-theoretic foundation for model selection and compression in deep learning, bridging theoretical guarantees with practical neural architecture design.

Addressing optimization challenges in neural network compressionBridging Kolmogorov complexity with deep learning theoryDeveloping optimal description length objectives for Transformers

Sequential Design with Posterior and Posterior Predictive Probabilities

Apr 01, 2025
LH
Luke Hagar
🏛️ McGill University | McGill University Health Centre

In Bayesian sequential trials, error rate evaluation relies on computationally expensive Monte Carlo simulations, hindering efficient optimization of sample size and decision thresholds. Method: This paper establishes, for the first time, analytical functional relationships between posterior and posterior predictive probabilities and sample size. Leveraging Bayesian decision theory and asymptotic analysis—combined with numerical fitting and error-rate inversion—the method enables precise error-rate assessment for any sample size using only two simulations, and rapidly identifies optimal design parameters. Contribution/Results: The approach drastically reduces computational cost while achieving error-rate control accuracy comparable to conventional simulation-based methods. In two real-world case studies, it attains exact error-rate calibration and accelerates design optimization by several orders of magnitude. This provides a scalable, verifiable, and highly efficient design paradigm for Bayesian adaptive trials.

Efficient error rate assessment for Bayesian sequential designsModeling probabilities as functions of sample sizeOptimal sample size determination using posterior probabilities

Automatic Parameter Selection for Non-Redundant Clustering

Dec 19, 2023
CL
Collin Leiber
🏛️ LMU Munich | University of Vienna

Subspace clustering in high-dimensional data often yields multiple semantically distinct subspaces, yet existing methods require manual specification of both the number of subspaces and the number of clusters within each—rendering them parameter-sensitive and poorly interpretable. This paper proposes an automatic, non-redundant multi-subspace clustering framework. First, it introduces the Minimum Description Length (MDL) principle to non-redundant clustering, enabling joint, adaptive inference of both the optimal number of subspaces and the cluster count per subspace. Second, it designs a split-merge-based greedy search strategy coupled with a subspace-level outlier encoding mechanism, allowing simultaneous outlier detection. Evaluated on multiple benchmark datasets, the method achieves competitive accuracy against state-of-the-art approaches while significantly improving parameter robustness, model interpretability, and practical applicability.

Automatically selects parameters for non-redundant clusteringDetects subspaces and clusters without user inputIdentifies outliers within each subspace efficiently

Latest Papers

What's happening recently
View more

This work addresses the challenge of risk prediction when acquiring true outcomes is prohibitively expensive, allowing labels for only a subset of samples. The authors propose a surrogate-assisted optimal sampling framework that, under a fixed annotation budget, leverages covariates, surrogate variables, and an initial estimator to construct a sampling strategy minimizing the expected out-of-sample cross-entropy loss. Coupled with an inverse probability weighted cross-entropy estimator for model training, this approach achieves—without requiring access to true responses during design—theoretically optimal sampling. It uniquely guarantees predictive optimality, robustness to surrogate misspecification, and stability in settings with rare outcomes. Both theoretical analysis and empirical experiments demonstrate that the method significantly outperforms existing approaches, particularly when the surrogate is imperfect or the event of interest is rare.

budget allocationmeasurement constraintsoptimal sampling

This work addresses the inefficiency of existing sampling-based inference methods, which struggle to effectively explore critical decision points during resampling. The authors propose a training-free inference optimization approach that identifies significant entropy spikes in the base model’s next-token prediction along the reasoning trajectory as proxies for key decision points. By integrating an entropy-guided cutoff selection strategy with the Metropolis-Hastings sampling algorithm, the method enables targeted resampling that substantially reduces mixing time complexity. Evaluated on multiple challenging benchmarks—including MATH500, HumanEval, GPQA Diamond, and AIME26—the approach consistently outperforms current baselines and even reinforcement learning–trained models, demonstrating strong training-free reasoning capabilities.

decision pointsmixing timepower distribution

While test-time scaling can enhance the reasoning performance of large language models, it incurs substantial computational overhead and latency. Existing adaptive sampling methods often rely on heuristic rules or distributional assumptions, making it challenging to efficiently balance accuracy and cost. This work formulates adaptive sampling as a Markov decision process and introduces the first lightweight reinforcement learning–based controller that operates solely on final-answer statistics, enabling training and deployment entirely on CPU. By incorporating Lagrangian relaxation, the approach achieves joint optimization under explicit budget constraints. Experiments demonstrate that the proposed method significantly improves the trade-off among accuracy, number of sampling rounds, and total sample count across multiple strong baselines, including ASC and ESC.

adaptive samplingcomputation costlarge language models

Hot Scholars

QL

Qian Li

Shenzhen Research Institute of Big Data
Theoretical computer sciencesArtificial Intelligence
MA

Matan Abudy

Tel-Aviv University
Computational Linguistics
SH

Shang-Hua Teng

University Professor of Computer Science and Math, USC
algorithmscomplexity theorygame theorynetwork science
RK

Roni Katzir

Tel Aviv University
linguisticscomputational linguisticsartificial intelligencelearnability