Score
Designs and runs benchmark suites, stress tests, and evaluation pipelines that compare and characterize mutual information estimators and estimators’ failure modes across regimes and constraints, covering nonparametric, discriminative, generative, zero-shot, and single‑pass estimators. Builds and analyzes algorithms and practical components that use MI objectives—including mutual information maximization, information‑theoretic optimization, adaptive or instance‑specific thresholding, and dataset‑level MI prediction—while identifying regime‑specific estimator winners, fundamental estimation barriers, and operational thresholds for downstream use.
Existing benchmarks for mutual information estimation are largely confined to low-dimensional, simplified distributions, limiting their ability to evaluate estimator performance on complex real-world data. This work proposes a unified benchmark framework grounded in copula theory, comprising two complementary test suites: the first systematically controls mutual information, dimensionality, and marginal complexity using synthetic data and flow-based models; the second integrates real images with controllable dependency structures, extending the classic same-pair paradigm. For the first time, this framework jointly encompasses both synthetic and real data while accounting for both dependency structure and marginal complexity. Systematic evaluation of diverse discriminative and generative estimators reveals that no single method dominates across all settings, and that fundamental limitations exist across different estimator families—limitations more effectively exposed by the newly designed tests.
Estimating mutual information (MI) in high-dimensional settings suffers from low efficiency, poor accuracy, and limited stability—especially when MI values span several orders of magnitude. To address this, we propose the first normalized flow-based MI estimator grounded in flow matching (FM). Unlike prevailing discriminative approaches, our method directly models the invertible density transformation between the joint and marginal distributions via normalizing flows, enabling end-to-end, differentiable MI estimation. By integrating flow matching into the MI estimation framework, we achieve both theoretical rigor—guaranteeing unbiased gradient estimation—and computational scalability. Experiments on multivariate benchmark tasks demonstrate that our estimator significantly outperforms baselines including InfoNCE and MINE in estimation accuracy, converges faster, incurs lower computational overhead, and exhibits strong robustness to both extremely small and large MI values.
In data-driven optimization, decision samples often exhibit optimistic bias relative to true performance due to the “optimizer’s curse.” To address this, we propose a first-order bias correction method that avoids re-optimization. We introduce the Optimizer’s Information Criterion (OIC), the first information-theoretic criterion tailored for decision selection in data-driven optimization—generalizing the Akaike Information Criterion (AIC) to encompass empirical models, parametric models, regularization, and contextual optimization. Leveraging asymptotic statistical analysis, we derive an analytical bias expression that explicitly captures the coupling between optimization and learning, eliminating the need for cross-validation. Evaluated on both synthetic and real-world datasets, our method achieves more accurate bias estimation and significantly lower computational overhead, while providing rigorous theoretical guarantees.
This work addresses the tightness of information-theoretic generalization error bounds with respect to sample size $n$, particularly the looseness of the individual-sample mutual information (ISMI) bound. To overcome the suboptimal $O(1/sqrt{n})$ convergence rate, we introduce, for the first time, an *excess risk assumption*, yielding a tight $O(1/n)$ fast-rate bound. Furthermore, we propose a novel generalization framework based on the $(eta,c)$-central condition, under which the mutual information term directly governs the convergence rate. We rigorously prove that this bound achieves the optimal $O(1/n)$ rate under standard assumptions. Empirical evaluation on canonical tasks—such as Gaussian mean estimation—demonstrates substantial improvements over existing information-theoretic bounds. The proposed framework thus bridges theoretical rigor with practical superiority, advancing both the tightness and applicability of information-theoretic generalization analysis.
This work addresses the challenge of estimating small failure probabilities under stochastic inputs in computationally expensive deterministic simulations. We propose a two-stage adaptive budget allocation framework: in Stage I, a Gaussian process surrogate is sequentially trained using a contour-localization strategy; in Stage II, remaining simulation budget is greedily allocated to critical regions—guided by classification entropy—to perform high-fidelity evaluations. A hybrid Monte Carlo estimator is then constructed by integrating surrogate predictions with observed high-fidelity responses. Our method introduces the first “exploration–exploitation decoupled” budget allocation paradigm, overcoming reliability limitations inherent in pure surrogate-based Monte Carlo and importance sampling. Experiments across multiple benchmark functions and an airfoil flow simulation demonstrate that the approach achieves significantly improved accuracy and robustness using only several hundred high-fidelity evaluations.
This work addresses the challenge that existing model evaluation methods often fail to reliably assess estimator quality in low-variance settings due to confounding between bias and variance or excessive sensitivity of statistical tests. To overcome this limitation, the authors propose a fault-tolerant evaluation framework that unifies bias and variance modeling through an adjustable tolerance parameter ε, enabling robust assessment of sample-efficient performance estimators within practically acceptable error margins. The framework integrates bias-variance analysis, fault-tolerant evaluation theory, and an adaptive ε-optimization algorithm, making it particularly well-suited for scenarios with low annotation costs. Experimental results demonstrate that the proposed approach provides a more comprehensive and reliable characterization of estimator behavior, significantly enhancing both the practical utility and stability of performance evaluation.
This work addresses key limitations of conventional mutual information (MI) estimators—poor generalization, computational inefficiency, and inability to quantify estimation uncertainty. We propose a fully data-driven neural estimator designed to overcome these challenges. Methodologically, we introduce an end-to-end differentiable neural architecture incorporating a 2D permutation-invariant attention mechanism to model joint distributions, employ quantile regression to produce calibrated uncertainty intervals, and leverage normalizing flows to synthesize a diverse, multimodal, multiscale meta-dataset for fully supervised training. Compared to classical and state-of-the-art neural MI estimators, our approach achieves significantly higher estimation accuracy across varying sample sizes and high-dimensional settings, accelerates inference by several orders of magnitude, yields more reliable confidence intervals, and natively integrates into end-to-end learning pipelines.
Mutual information (MI) estimation is fundamental in data science, yet existing approaches struggle to reconcile model flexibility with statistical interpretability: neural estimators require substantial data, while classical models (e.g., Gaussian copulas) fail to capture complex, high-order dependencies. This paper introduces the first neural MI estimation framework grounded in vector vine structures—marking the first integration of vector vine theory into deep learning–based estimation. Methodologically, we employ neural networks to parameterize marginal transformations and leverage vector vines to explicitly model intricate multivariate dependence structures; training is performed end-to-end via a variational lower bound and density-ratio estimation. Evaluated on synthetic benchmarks and multimodal real-world datasets, our approach achieves significant gains in estimation accuracy, robustness, and generalization—effectively balancing expressive power and statistical interpretability.
This work investigates the fundamental performance limits of learning and estimation tasks within an information-theoretic framework, independent of the computational capabilities of specific algorithms. By integrating tools from information theory and statistical learning theory—including metric entropy, VC dimension, Rademacher complexity, mutual information, and relative entropy—it systematically derives multiple upper bounds on generalization error. Simultaneously, leveraging Fano’s inequality together with covering and packing numbers, the study establishes information-theoretic lower bounds on minimax risk. The analysis unifies two complementary paradigms: one grounded in the geometric structure of metric spaces and the other based on information-theoretic measures. This synthesis yields a rigorous and broadly applicable theoretical framework for characterizing the optimal performance boundaries inherent to learning and estimation problems.
This work employs information-theoretic tools to understand and optimize the training dynamics of statistical learning models, with a particular focus on generative models. By integrating key concepts such as f-divergence, Fisher divergence, and the evidence lower bound (ELBO), it establishes a unified framework that systematically encompasses a broad spectrum of methods—from linear regression to diffusion models. Notably, the paper provides a more explicit and systematic derivation of generative diffusion models than existing treatments in the literature. Beyond offering deeper information-theoretic insights into model training mechanisms, this study also delivers a pedagogically structured exposition well-suited for teaching and self-study, thereby presenting a cohesive information-theoretic perspective across multiple mainstream modeling paradigms.