Score
Designs, trains and deploys models and pipelines that assign each input to one of two classes, including choices of loss functions, algorithms, thresholding/quantization, and training procedures for binary decision tasks. Evaluates and analyzes these models using binary-specific metrics and practices (ROC/AUC, precision/recall, FPR/FNR, calibration), frames tasks as two-class problems, optimizes thresholds for operational trade-offs, and assesses robustness and generalization across dataset shifts.
Conventional argmax decision rules in multi-class classification are non-differentiable and lack a principled, threshold-based optimization mechanism analogous to that in binary classification. Method: This paper proposes a simplex-geometric posterior optimization framework with multi-dimensional thresholds, reframing classification as a thresholding operation on the probability simplex—bypassing softmax’s probabilistic interpretation and enabling plug-and-play performance enhancement for arbitrary pre-trained models. Contribution/Results: We introduce ROC-cloud analysis and DFP (Distance From Point) scoring, yielding a more consistent, differentiable, and interpretable evaluation paradigm than One-vs-Rest. Extensive experiments across diverse network architectures and benchmark datasets demonstrate significant improvements in both accuracy and robustness, validating the efficacy of multi-dimensional threshold tuning and the generality of the proposed evaluation framework.
This work investigates the geometric foundations of ROC and PR curves in binary classification, aiming to unify the understanding of curve morphology and classifier behavior through a geometric lens. Methodologically, it introduces the composite function (G = F_p circ F_n^{-1}) as a core modeling framework—where (F_p) and (F_n) denote the CDFs of positive and negative class score distributions—and rigorously establishes a geometric mapping between ROC/PR curve shapes and the underlying distributional geometry. It reveals that (G) quantifies inter-class leakage and admits interpretation via KL divergence. Furthermore, it derives geometric criteria for classifier dominance and interpretability grounded in differential geometry, statistical inference, and CDF transformation theory. The contributions include: (i) a principled, geometrically interpretable framework for threshold selection; (ii) robust, distribution-agnostic tools for classifier comparison; and (iii) enhanced reliability and adaptability in cost-sensitive deployment—particularly under class imbalance and distributional overlap.
This work addresses the problem of shortcut learning in binary black-box classification models caused by dataset bias. It proposes a novel post-hoc analysis framework that integrates interventional and observational perspectives, introducing linear mixed-effects models—used here for the first time—to diagnose bias in black-box classifiers. By decomposing the influence of training and test data on model scores, the method moves beyond conventional error-rate metrics to uncover the risk of models relying on spurious correlations. The approach effectively identifies and quantifies the impact of data bias on decision-making in voice anti-spoofing and speaker verification tasks, offering a new pathway toward building reliable and interpretable AI systems.
This paper addresses the problem that conventional calibration evaluation of deep learning models is vulnerable to spurious recalibration—i.e., trivial post-hoc adjustments that improve calibration metrics without enhancing generalization. To tackle this, we propose a novel joint evaluation paradigm integrating calibration and generalization. First, we derive a Bregman-divergence-based decomposition of calibration error, establishing the first theoretical connection between calibration metrics and generalization objectives (e.g., negative log-likelihood). Second, we design a new reliability diagram that jointly visualizes calibration bias and estimated generalization error. Third, we characterize multiple “pseudo-optimal” calibration phenomena and provide theoretically grounded, detectable criteria for identifying trivial recalibration. Experiments on standard benchmarks demonstrate that our approach significantly improves model diagnostic capability, yielding a more reliable and interpretable evaluation framework for calibration research.
The widespread deployment of machine learning (ML) in decision-making systems introduces significant fairness risks—particularly concerning the handling of sensitive attributes and the protection of minority groups—while software engineering lacks a systematic, lifecycle-oriented framework for fairness engineering practices. Method: We conduct a systematic mapping study (SMS) combined with a comprehensive literature review to analyze fairness-related practices across the ML development lifecycle. Contribution/Results: We propose the first software engineering–centric fairness practice taxonomy, comprising 28 structured, actionable practices explicitly mapped to data preprocessing, modeling, and deployment stages. Each practice is annotated with its corresponding ML lifecycle phase and contextual applicability, thereby bridging the gap between fairness research and industrial implementation. This taxonomy serves as an integrable, operational guide for researchers and practitioners, enhancing the reliability, accountability, and trustworthiness of ML systems.
This work addresses the potential nonexistence of a global optimum in linear ensembles of multiple binary classifiers by proposing a theoretical framework grounded in truth-table logical structuring and equivalence class partitioning, which establishes sufficient conditions for the existence of a convexified empirical risk minimizer. By introducing a multidimensional generalization of classification-calibrated loss functions and the notion of φ-frontiers, the study analyzes solution stability in relation to data quality. Under exponential (Boost) and logistic (Logit) losses, the authors derive, for the first time, explicit closed-form expressions for the optimal ensemble weights and fully characterize all solution regimes in the three-classifier setting. This approach circumvents iterative optimization, thereby substantially enhancing both the interpretability and computational efficiency of ensemble models.
This work addresses the challenge of prediction score distribution drift caused by frequent model retraining in security applications, which undermines downstream systems’ reliance on consistent false positive rates (FPR). To resolve this, the authors propose a novel binary classifier calibration method that operates across the full FPR spectrum, departing from conventional probability-based calibration paradigms. Their approach enables precise control over output scores at any desired FPR threshold, ensuring semantic consistency of FPR across different model versions. Built upon existing calibration primitives, the method directly optimizes FPR contracts rather than class probabilities, making it suitable for large-scale production deployment. Experiments demonstrate that within the 0.01%–10% FPR range, the relative FPR error remains below 7.2%, with a calibration artifact size under 200 KB, and the solution scales efficiently from thousands to millions of samples.
This study addresses the limitations of evaluating multiclass classifiers using single performance metrics, which often leads to misleading conclusions. To overcome this, the work proposes a multidimensional evaluation paradigm that leverages the PyCM library to construct a comprehensive analytical framework, enabling systematic comparison of classifier performance across a diverse set of evaluation metrics. Through two case studies, the research uncovers nuanced performance trade-offs that conventional metrics fail to capture, thereby demonstrating the necessity and effectiveness of multidimensional assessment in model selection and optimization. The findings further highlight the unique value of PyCM in facilitating thorough and precise evaluation of multiclass classification systems.
This work addresses the lack of systematic methodologies in model optimization, which often relies on heuristic choices and struggles to accommodate diverse deployment constraints. It formalizes model compression and acceleration as a constraint-aware multi-objective engineering decision problem, establishing a unified and actionable framework grounded in five key dimensions: data availability, latency, memory footprint, accuracy tolerance, and retraining budget. By integrating techniques such as quantization, pruning, knowledge distillation, parameter-efficient fine-tuning (PEFT), and inference optimization, the study proposes tailored optimization pipelines for four representative industrial scenarios, delivering a reproducible and quantifiable guide for technology selection.
本文探讨了在数据不完美条件下的机器学习挑战,并通过信息损失、经验风险偏差等机制组织代表性方法,如重建生成、再平衡与表示校准等。