Score
A family of convex-function–derived discrepancy measures used to quantify geometry and directional differences between predictions and targets; used to define impurity/impurity-gain for trees, construct geometry-preserving neural scoring heads, and characterize population-level limits of calibration procedures.
This work proposes a unified decision tree framework grounded in Bregman divergences, addressing the limitation of traditional methods—such as CART—that rely on ad hoc and isolated impurity criteria lacking a cohesive theoretical foundation. By systematically integrating principles from convex analysis and information geometry into the tree construction process, the framework derives a general class of impurity measures applicable to a broad range of statistical models, including exponential family distributions. The approach not only subsumes numerous existing loss functions but also leverages strong convexity and smoothness properties to establish a family of generalized decision trees with rigorous theoretical guarantees, such as stability and consistency. This significantly enhances the model’s adaptability to diverse data distributions and underlying geometric structures.
This work investigates the geometric foundations of ROC and PR curves in binary classification, aiming to unify the understanding of curve morphology and classifier behavior through a geometric lens. Methodologically, it introduces the composite function (G = F_p circ F_n^{-1}) as a core modeling framework—where (F_p) and (F_n) denote the CDFs of positive and negative class score distributions—and rigorously establishes a geometric mapping between ROC/PR curve shapes and the underlying distributional geometry. It reveals that (G) quantifies inter-class leakage and admits interpretation via KL divergence. Furthermore, it derives geometric criteria for classifier dominance and interpretability grounded in differential geometry, statistical inference, and CDF transformation theory. The contributions include: (i) a principled, geometrically interpretable framework for threshold selection; (ii) robust, distribution-agnostic tools for classifier comparison; and (iii) enhanced reliability and adaptability in cost-sensitive deployment—particularly under class imbalance and distributional overlap.
When machine learning models are employed as measurement instruments, it remains unclear whether their outputs genuinely reflect stable and consistent latent constructs beyond merely achieving predictive performance. This work formally introduces the concept of “learned measurement functions” and proposes “measurement stability” as a distinct evaluation criterion. Through theoretical analysis and empirical case studies, we demonstrate that conventional metrics—such as generalization error, calibration, and robustness—do not guarantee measurement consistency. Our findings reveal that models with comparable predictive accuracy can implement systematically inequivalent measurement functions, and that these discrepancies become pronounced under distributional shifts, thereby exposing critical limitations in current evaluation frameworks.
This work addresses the high $L_\infty$ star discrepancy of three-dimensional Kronecker point sets at fixed sizes by proposing a parameter optimization approach based on the algorithm configuration framework irace. By automatically tuning the two generating parameters of Kronecker sequences, the method systematically searches for optimal configurations across various point set sizes, particularly those with at least 500 points. This study marks the first application of irace to the construction of Kronecker point sets and achieves significantly better performance than existing methods across multiple continuous size intervals, establishing new state-of-the-art records for $L_\infty$ star discrepancy. The resulting point sets exhibit enhanced uniformity, offering improved suitability for applications such as experimental design, Bayesian optimization, and quasi-Monte Carlo integration.
Neural network probability outputs in multiclass classification are often miscalibrated, and existing post-hoc calibration methods lack theoretical foundations. Method: We propose ALR (Additive Log-Ratio) calibration—a geometrically principled, information-geometric framework—extending ALR calibration to multiclass settings and naturally generalizing Platt scaling. By leveraging the Fisher–Rao metric on the probability simplex, we construct an interpretable calibration mapping and integrate a geometric reliability score with a distance-based neutral-zone decision rule to provably detect erroneous predictions. Contribution/Results: Theoretically, we establish calibration consistency and derive a concentration bound for reliability using M-estimation theory, sub-Gaussian inequalities, and the Bhattacharyya coefficient. Empirically, on a viral classification task, our two-stage method captures 72.5% of prediction errors at a 34.5% sample abstention rate, providing a statistically grounded, actionable mechanism for uncertainty-aware prediction.
This study addresses the inconsistency of existing prediction evaluation metrics—such as ABC and Gini—with the principle of mean consistency, stemming from their reliance on predicted values for weighting, which can lead to erroneous model selection. Building upon Bregman divergences, the authors develop a mean-consistent loss framework, rederive the Murphy decomposition to disentangle prediction error into calibration and discrimination components, and establish a theoretical link between these components and Lorenz-curve-based metrics. They propose a new metric, ABC², to enhance sensitivity to mean calibration, and demonstrate that ABC, ABC², and Gini all violate mean consistency due to prediction-dependent weighting. Furthermore, they prove the equivalence between the number of crossings in Lorenz and Murphy curves and, under a single-crossing condition, provide a weak dominance criterion for predictive superiority, offering both theoretical grounding and practical guidelines for reliable model evaluation.
To address the limited discriminative power of Maximum Mean Discrepancy (MMD) in goodness-of-fit testing, this paper proposes a spectral-filtering-based regularized kernel discrepancy framework. The method constructs flexible test statistics via integral operators, relaxing stringent assumptions on kernels and filter functions inherent in prior approaches, and achieves rigorous Type-I error control with improved statistical power in non-asymptotic settings. Theoretically, the proposed test constitutes a natural generalization of existing MMD-based tests, offering enhanced detection sensitivity and broader theoretical applicability. Empirical evaluations demonstrate that it matches or outperforms state-of-the-art methods across diverse scenarios—including multivariate, high-dimensional, and small-sample settings—exhibiting strong practical adaptability and competitive performance.
This work addresses the limitations of traditional generalization analyses, which rely on the often unverifiable assumption of independent and identically distributed (i.i.d.) data and thus struggle to accurately characterize model performance on unseen data. The paper proposes a deterministic generalization analysis framework that dispenses with any prior probabilistic assumptions. By examining the sensitivity of optimization solutions to data perturbations, it decomposes the generalization error into geometric and probabilistic components, achieving their first-ever decoupling. The framework expresses generalization bounds via a variational principle, leveraging deterministic perturbation analysis and optimization sensitivity theory to capture the discrepancy between in-sample and out-of-sample performance. Error terms are evaluated through posterior statistical hypotheses, enabling the recovery of conventional high-probability or expected generalization guarantees—all without requiring distributional assumptions.
This work proposes a unified geometric framework for the joint analysis of prediction, interpretability, robustness, and generalization in tree ensemble models. By leveraging node-weighted feature mappings and path-isometric embeddings, the framework constructs a non-diagonal Gram kernel representation in a squared Euclidean space, thereby realizing a metric embedding of tree ensembles. For the first time, it integrates exact additive attributions, deterministic Lipschitz robustness radii (based on the KPP metric), and Rademacher generalization bounds within a single kernel structure. Under honesty and cross-fitting conditions, the authors derive a unified risk bound applicable to both regression and classification. The study further establishes the existence of fast convergence rates and formulates related open problems.
This work addresses the potential nonexistence of a global optimum in linear ensembles of multiple binary classifiers by proposing a theoretical framework grounded in truth-table logical structuring and equivalence class partitioning, which establishes sufficient conditions for the existence of a convexified empirical risk minimizer. By introducing a multidimensional generalization of classification-calibrated loss functions and the notion of φ-frontiers, the study analyzes solution stability in relation to data quality. Under exponential (Boost) and logistic (Logit) losses, the authors derive, for the first time, explicit closed-form expressions for the optimal ensemble weights and fully characterize all solution regimes in the three-classifier setting. This approach circumvents iterative optimization, thereby substantially enhancing both the interpretability and computational efficiency of ensemble models.