Score
Design, implement, and evaluate algorithms and estimators that compute, estimate, or compare entropy and related information measures (e.g., Shannon entropy, spectral entropy, entropy weighting) from probability distributions, empirical samples, model outputs, or signal spectra. This includes deriving and applying entropy inequalities (including the entropy power inequality), building maximum-entropy models and entropy-minimization procedures, and using entropy-based cutoffs or uncertainty metrics for analysis and decision-making.
Accurate estimation of entropy, mutual information, and conditional mutual information in software engineering is often hindered by high computational cost and long runtime. This paper systematically evaluates 18 bias-corrected entropy estimators across varying sample sizes and domain cardinalities. Through large-scale simulations of random joint distributions—complemented by rigorous statistical bias analysis and quantitative convergence assessment—we identify, for the first time, that the Chao–Shen and Chao–Wang–Jost estimators consistently exhibit rapid convergence and strong robustness across all entropy measures. Crucially, they achieve superior accuracy and faster convergence under low-sample-size conditions. Our findings yield a lightweight, reliable, and plug-and-play entropy estimation framework, directly applicable to software confidentiality analysis, test adequacy assessment, and machine learning feature selection.
Reliable estimation of Shannon entropy from small samples suffers from systematic underestimation, particularly when the sample size is smaller than the number of possible outcomes. To address this, we propose a novel discrete entropy estimator that jointly models sample-space partitioning, missing-mass estimation, and unseen-outcome count estimation—leveraging entropy’s decomposability, interpolation-based estimation, empirical frequency correction, and probabilistic modeling of unseen events for synergistic bias correction. The method substantially mitigates negative bias and significantly outperforms classical estimators—including maximum likelihood estimation (MLE) and jackknife—in the undersampled regime. Its performance matches that of state-of-the-art approaches while exhibiting superior robustness and higher accuracy across diverse distributions and sampling conditions.
Classical maximum entropy principle (MEP) relies on the system independence assumption in the Shore–Johnson axioms, which frequently fails in strongly correlated systems (e.g., economic or ecological networks), leading to systematic biases in Shannon-entropy-based inference. Method: We demonstrate that the Uffink–Jizba–Korbel (UJK) one-parameter generalized entropy family relaxes this assumption, offering a more robust entropy selection criterion for non-independent systems. By reformulating the Shore–Johnson axiomatization, we precisely delineate the domain of applicability for UJK entropies and establish a reproducible, transparent framework for entropy function selection and reporting. Contribution/Results: Empirical validation in economics (market interdependence modeling) and ecology (inference of species interactions) shows substantial improvements in distribution reconstruction accuracy and interpretability. This work is the first to systematically bridge foundational axiomatic principles with practical implementation guidelines, thereby advancing the reliable application of MEP in complex systems.
This paper addresses the fragmentation and lack of a unified theoretical foundation for entropy measures in data analysis and machine learning. We propose the first comprehensive, multi-source entropy theory taxonomy and generic framework tailored to data science. Grounded in the Shannon–Khinchin axioms, it unifies over ten entropy families—including Shannon, Rényi, Tsallis, Kolmogorov–Sinai, and von Neumann entropies—bridging perspectives from information theory, statistical learning, dynamical systems, and quantum probability. We systematically catalog over one hundred entropy-driven algorithms and empirically demonstrate that our framework significantly enhances robustness and interpretability in high-dimensional dimensionality reduction, time-series pattern recognition, and uncertainty quantification. Our core contributions are threefold: (1) a formal axiomatized generic representation of entropy; (2) a systematic methodology for entropy-based learning; and (3) a unified theoretical foundation for feature selection, clustering, anomaly detection, and model interpretation.
This work investigates the fundamental performance limits of learning and estimation tasks within an information-theoretic framework, independent of the computational capabilities of specific algorithms. By integrating tools from information theory and statistical learning theory—including metric entropy, VC dimension, Rademacher complexity, mutual information, and relative entropy—it systematically derives multiple upper bounds on generalization error. Simultaneously, leveraging Fano’s inequality together with covering and packing numbers, the study establishes information-theoretic lower bounds on minimax risk. The analysis unifies two complementary paradigms: one grounded in the geometric structure of metric spaces and the other based on information-theoretic measures. This synthesis yields a rigorous and broadly applicable theoretical framework for characterizing the optimal performance boundaries inherent to learning and estimation problems.
This work addresses the lack of intuition in traditional derivations of exponential family distributions, which often obscure their information-theoretic and physical foundations in pedagogical contexts. By leveraging the principle of maximum entropy and requiring only elementary notions of entropy, the paper presents a concise and self-contained derivation that avoids complex constrained optimization. The core contribution demonstrates that, under constraints fixing the expected values of sufficient statistics, exponential family distributions uniquely maximize relative entropy with respect to a general base measure, and Shannon entropy in the special case of a uniform base measure. This approach reveals the fundamental connection between maximum entropy and exponential families from minimal assumptions, substantially streamlining the didactic exposition and fostering deeper integration of statistical theory with physical reasoning.
This work investigates the sample complexity of estimating Rényi entropy and min-entropy under high-dimensional discrete distributions. By establishing constructive upper bounds and information-theoretic lower bounds, it provides the first tight characterization of min-entropy estimation with sample complexity Θ(k log k), correcting prior erroneous claims. For integer-order Rényi entropy, matching upper and lower bounds are derived, revealing the critical role of the order α in determining sample complexity. The analysis leverages an unbiased falling-factorial estimator for α-wise collisions, a concentration inequality based on dyadic interval partitioning, and a hidden-heavy-coordinate construction. The results show that when 1.001 ≤ α ≤ c₀ log k, the sample complexity is Θ_{c₀}(αk^{1−1/α}), while for higher orders it becomes Θ_ε(k log k).
This work addresses the lack of robustness in ordinary least squares when sparse, large-magnitude outliers are present. The authors propose a maximum-entropy-based weighted least squares framework, wherein weights are interpreted as a discrete probability distribution and determined by maximizing Shannon entropy subject to a mean squared error constraint. This approach automatically downweights outliers while preserving fidelity to inliers. Innovatively treating the mean squared error as a tunable control parameter, the method leverages the principle of maximum entropy to select the least biased weight distribution. The paper establishes the theoretical existence of a locally unique, globally continuable smooth solution branch emanating from the ordinary least squares solution. Further analysis reveals that, in the zero-error limit, the method automatically identifies the largest subset of data consistent with the underlying linear model. Numerical experiments confirm its superior robustness.
This study investigates the sample complexity required to distinguish two probability distributions based on their Jensen–Shannon divergence (JSD). Focusing on independent and identically distributed samples, the authors analyze the logarithmic likelihood ratio classifier and the majority vote classifier under a fixed JSD. By leveraging tools from information theory and statistical learning theory, they establish distinct scaling laws for the sample complexity of the two classifiers: it grows as $1/\text{JSD}$ for the likelihood ratio classifier and as $1/\text{JSD}^2$ for the majority vote classifier. These findings provide an operational statistical interpretation of JSD and, for the first time, quantify its precise relationship with the distinguishability of probability distributions.
This study addresses the challenge of achieving both efficiency and robustness in parameter estimation with partially observed data under the missing-at-random (MAR) mechanism. The authors propose a unified inference framework based on generalized entropy calibration weighting. By solving a convex entropy minimization problem subject to balancing and debiasing constraints, the method constructs weights that integrate data-adaptive calibration functions, flexible machine learning predictors, cross-fitting, and propensity score modeling. This approach subsumes inverse probability weighting (IPW) and augmented IPW (AIPW) as special cases, enjoys double robustness, and attains the semiparametric efficiency bound when both outcome and propensity models are correctly specified. Simulation and empirical analyses demonstrate that the proposed estimator outperforms existing methods in terms of statistical efficiency and numerical stability, particularly exhibiting substantial gains over standard AIPW when the outcome model is misspecified.