Score
Designs and analyzes mathematical descriptions of the relationship between (typically convex) loss-generating functions and the Bregman divergences they induce, and applies Bregman divergence theory to derive quantitative statements about estimators and algorithms. Uses Bregman divergence to relate losses to error measures, characterize target function spaces, and prove guarantees such as population-level limits, convergence rates, calibration, and other algorithmic properties.
This work proposes a unified decision tree framework grounded in Bregman divergences, addressing the limitation of traditional methods—such as CART—that rely on ad hoc and isolated impurity criteria lacking a cohesive theoretical foundation. By systematically integrating principles from convex analysis and information geometry into the tree construction process, the framework derives a general class of impurity measures applicable to a broad range of statistical models, including exponential family distributions. The approach not only subsumes numerous existing loss functions but also leverages strong convexity and smoothness properties to establish a family of generalized decision trees with rigorous theoretical guarantees, such as stability and consistency. This significantly enhances the model’s adaptability to diverse data distributions and underlying geometric structures.
This paper addresses the equivalence between two definitions of weighted vector information—based on variability and non-uniformity—and establishes that Bregman divergence is the unique class of divergences ensuring their strict identity. Using tools from convex analysis, information geometry, and functional equations, we provide the first rigorous characterization showing that information-equivalence fully determines the functional form of the divergence. This yields a necessary and sufficient condition linking information-theoretic equivalence to divergence structure, thereby proving the uniqueness and foundational role of Bregman divergences within divergence theory. As a key contribution, we derive a novel axiomatic characterization of Bregman divergences, offering an essential theoretical criterion for divergence selection in optimization, statistics, and information theory.
This work addresses the minimal parameter count required for Lipschitz-robust interpolation by overparameterized models. Extending the Bubeck–Sellke lower bound—originally established for squared loss and scalar responses—to general Bregman divergences (e.g., cross-entropy, squared error) and vector-valued outputs, the authors reformulate the proof within a bias–variance decomposition framework, circumventing reliance on Rademacher complexity. Instead, they directly leverage properties of Bregman divergences, concentration inequalities, and Lipschitz constraints. Their key contribution is the first unified robustness law: for any Bregman loss, $d$-dimensional inputs, and $m$-dimensional outputs, achieving Lipschitz interpolation necessitates $Omega(n + d)$ parameters, where $n$ is the sample size. This lower bound reveals the fundamental necessity of overparameterization in generalized loss settings and multi-output learning, unifying and generalizing prior results on robust interpolation.
Classical bias–variance decomposition is restricted to squared error loss, limiting its applicability in statistical learning where non-quadratic losses (e.g., log loss, exponential loss) are common. Method: Leveraging convex analysis and statistical decision theory, we generalize the decomposition to arbitrary Bregman divergences as prediction errors, deriving a rigorous formula for maximum likelihood estimators under exponential family distributions and specifying precise conditions for its validity. Contribution/Results: Our framework unifies previously fragmented results for specific losses, filling a fundamental theoretical gap. It enhances interpretability and pedagogical utility of the bias–variance trade-off, and provides a principled, general-purpose tool for model diagnostics and generalization analysis under non-squared error settings—enabling coherent error decomposition across diverse loss functions grounded in information geometry.
The original kernelized Kullback–Leibler (KL) divergence is ill-defined when the supports of compared distributions are disjoint—a fundamental limitation. To address this, the paper proposes a Tikhonov-regularized kernel KL divergence, constructed via covariance operator embeddings in a reproducing kernel Hilbert space (RKHS). This metric is well-defined for arbitrary probability distributions—including discrete, continuous, and mutually singular ones—and provides theoretical guarantees: a bias bound relative to the true KL divergence, finite-sample convergence rates, and a closed-form solution for discrete distributions. Furthermore, the authors formulate a Wasserstein gradient flow optimization framework for the proposed divergence, ensuring theoretical convergence, and design an efficient algorithm applicable to discrete structures such as point clouds. Experiments on point cloud transport tasks demonstrate that the method outperforms existing kernelized and Wasserstein-based approaches, achieving superior stability and robustness.
This work addresses the limited integration of existing functional Bregman divergences with kernel methods and reproducing kernel Hilbert spaces (RKHS), which has hindered their applicability in modern machine learning. The paper presents the first systematic incorporation of the RKHS framework into functional Bregman divergences, leveraging the Riesz representation theorem and self-dual pairings to simplify their structure. Building upon kernel mean embeddings, the authors derive a computationally efficient form of the divergence. This approach not only establishes a theoretical bridge between Bregman geometry and kernel methods but also unifies existing techniques such as maximum mean discrepancy (MMD) within a common framework. Empirical evaluations demonstrate that the proposed divergence achieves strong performance in tasks including clustering, robust estimation, and generative modeling.
This work addresses the challenge of achieving calibeating—simultaneously ensuring prediction calibration and low regret—under a broad family of proper loss functions, including α-Tsallis losses and log loss. By adopting a Bregman divergence perspective, we develop a unified framework that generalizes calibeating to arbitrary proper losses for the first time. Within this framework, we introduce a “Be The Regularized Leader” algorithm together with a novel regret identity. Our approach substantially weakens the dependence on dimensionality in U-calibration guarantees and yields logarithmic regret bounds for the Tsallis loss family, improving upon existing results. This provides a powerful new tool for balancing calibration and regret control in online learning settings.
This work investigates the fundamental performance limits of learning and estimation tasks within an information-theoretic framework, independent of the computational capabilities of specific algorithms. By integrating tools from information theory and statistical learning theory—including metric entropy, VC dimension, Rademacher complexity, mutual information, and relative entropy—it systematically derives multiple upper bounds on generalization error. Simultaneously, leveraging Fano’s inequality together with covering and packing numbers, the study establishes information-theoretic lower bounds on minimax risk. The analysis unifies two complementary paradigms: one grounded in the geometric structure of metric spaces and the other based on information-theoretic measures. This synthesis yields a rigorous and broadly applicable theoretical framework for characterizing the optimal performance boundaries inherent to learning and estimation problems.
This work addresses the limitations of conventional regression loss functions, which are typically based on absolute error and thus ill-suited for tasks involving multiplicative noise or where relative error is of primary concern. For the first time, the paper systematically investigates ratio-based loss functions defined in terms of the quotient between predicted and true values. Through rigorous mathematical analysis—agnostic to any specific learning algorithm—it examines fundamental properties such as continuity, Lipschitz continuity, convexity, and differentiability. The study fills a critical gap in the theoretical understanding of this class of losses, introduces several novel loss formulations, and establishes a general analytical framework that lays the groundwork for future research on consistency, learning rates, and algorithmic stability.
This study addresses the unclear impact of divergence selection on the performance and optimization efficacy of the Shampoo preconditioner. To this end, it constructs a unified theoretical framework based on Bregman divergences, integrating spectral analysis with GPT-2 pretraining experiments to systematically investigate how different divergences influence Kronecker approximations and finite-sample errors. The findings reveal that specific divergences can effectively compensate for the underestimation of second-order moments, offering a novel perspective for understanding Shampoo. Furthermore, this work elucidates the underlying mechanisms driving behavioral differences among various Shampoo variants, providing principled theoretical guidance for the design and refinement of second-order optimizers.