Score
Designs and performs analyses and visualizations of classification confusion matrices to compute and interpret class-wise error rates and misclassification patterns. Builds metrics and comparisons to identify systematic errors, ambiguous example groups, and differences in error distributions across models or baselines.
Confusion matrices are often confounded by both inter-class similarity and data distribution bias, making it difficult to disentangle their independent contributions to misclassification. To address this, we propose a doubly stochastic normalization method that achieves joint row- and column-wise normalization via iterative proportional fitting. This is the first approach to establish an explicit geometric correspondence between the normalized confusion matrix and the model’s class representation space. Our method effectively decouples distributional bias from intrinsic class confusion, thereby accurately recovering the underlying class similarity structure. Experiments demonstrate that the proposed normalization significantly improves the accuracy and interpretability of error pattern diagnosis. It enables fine-grained analysis of classifier behavior—distinguishing systematic biases due to imbalanced sampling from fundamental ambiguities arising from semantic or feature-space proximity—and provides a principled tool for classifier evaluation, debugging, and optimization.
Existing classification performance metrics suffer from inconsistent scales and high sensitivity to class imbalance, rendering cross-dataset evaluation incomparable and difficult to interpret. To address this, we propose the Outperformance Score (OS), a unified normalization framework that maps any confusion-matrix-based metric onto the [0,1] interval. OS is defined as the percentile rank of an observed performance value within a reference empirical distribution induced by the class imbalance ratio. This constitutes the first general-purpose standardization framework that endows diverse CMBCP metrics with consistent semantics and comparable scales. Crucially, OS eliminates reliance on fixed decision thresholds or parametric distributional assumptions, enabling robust evaluation under dynamically varying imbalance ratios. Extensive validation across real-world datasets from healthcare, finance, and natural language processing domains demonstrates that OS significantly enhances reliability and consistency in both inter-metric and cross-dataset comparisons, while supporting plug-and-play integration.
Current CVML models struggle to localize and attribute failures at the subgroup level without labeled data. Method: We propose the first interactive error analysis framework integrating large language model (LLM) semantic understanding with visual analytics—requiring no manual annotations. It leverages CLIP-based semantic embeddings to cluster misclassified images into semantically coherent subgroups, then employs GPT-4 to generate interpretable, hypothesis-driven explanations. The framework further supports concept-guided interactive validation and comparative analysis. Contribution/Results: Evaluated on image classification, object detection, and semantic segmentation, our approach significantly improves both the efficiency and depth of error attribution—enabling an explainable shift from “where errors occur” to “why they occur.” Domain experts have validated its effectiveness and practical utility.
To address the delayed performance monitoring caused by the absence of ground-truth labels post-deployment, this paper proposes the first label-free, multi-metric online estimation framework for binary classification tasks, supporting accuracy, precision, recall, and F1-score—metrics derived from the confusion matrix. Methodologically, it is the first to systematically generalize label-free performance estimation to *any* metric definable via the confusion matrix. By calibrating model confidence scores to model predictive probability distributions and leveraging stochastic variable propagation with Monte Carlo inference, the approach enables full probabilistic modeling of confusion matrix elements, yielding theoretically grounded point estimates and statistically valid confidence intervals. Experiments across multiple benchmark datasets demonstrate that the proposed method significantly outperforms existing baselines, achieving lower estimation errors across all metrics while maintaining nominal coverage rates for its confidence intervals.
Conventional binary classifier evaluation relies on single metrics, failing to holistically address diverse application scenarios. Method: This paper proposes a unified multi-perspective analysis framework based on Tile visualization, which innovatively maps infinite-dimensional ranking scores onto a two-dimensional Tile plot and geometrically models classifier behavior in ROC space—enabling performance comparison under arbitrary metric combinations. The framework supports four user categories (theoretical analysis, algorithm design, benchmarking, and application development) via customizable preference modeling and role-adapted “flavor” interpretations. Results: Empirical evaluation across 74 state-of-the-art semantic segmentation models demonstrates that a single Tile plot comprehensively captures model performance differences, significantly improving cross-task evaluation consistency, interpretability, and practical utility.
This study addresses the pervasive issue of data errors in real-world databases—such as missing values, redundancy, statistical biases, and outliers—which significantly degrade downstream analytical and machine learning performance. Recognizing that existing taxonomies are incomplete and terminology inconsistent, this work presents the first unified framework that integrates traditional data errors with statistically oriented inaccuracies critical in the AI era. It proposes a non-overlapping tripartite classification structure—comprising missing, erroneous, and redundant data—and systematically constructs a comprehensive catalog of 35 distinct error types. Through formal definitions, illustrative examples, and a thorough literature review, the paper establishes standardized terminology and precise characterizations, thereby offering a clear, rigorous theoretical foundation and practical toolkit for data quality assessment and cleaning.
Current evaluation practices for supervised learning models are often misleading due to an overreliance on single aggregate metrics, which neglect the alignment among data characteristics, task objectives, and real-world application contexts. This work reframes model evaluation as a context-dependent, decision-oriented process and systematically investigates—through controlled experiments—the impact of dataset properties, validation strategies, class imbalance, and asymmetric error costs on evaluation outcomes. Leveraging diverse benchmark datasets, multiple validation protocols, and multidimensional performance measures, the study uncovers common pitfalls such as the accuracy paradox, data leakage, and metric misuse. It proposes a structured evaluation framework explicitly aligned with operational goals, offering principled guidance for developing more robust, reliable, and trustworthy supervised learning systems.
Real-world datasets often suffer from multiple issues, including label noise, feature corruption, and spurious correlations, yet existing methods struggle to simultaneously identify erroneous samples and their specific error types with high precision. This work proposes DeMix, a novel framework that, for the first time, leverages influence vectors to characterize how individual training samples affect predictions on a validation set. By formulating data debugging as a multi-label classification task and incorporating intervention-based learning to extract invariant diagnostic criteria for each error type, DeMix enables joint identification of both corrupted samples and their underlying error categories. Evaluated across 11 benchmark tasks, DeMix improves the F1 score for data debugging by 22.61% on average and boosts downstream model performance by 9.32% after data repair, substantially outperforming current state-of-the-art approaches.
This work addresses the challenge of attributing misclassifications and evaluating robustness in black-box classifiers by proposing an explainability-aware optimization framework. The approach integrates L₀ sparsity regularization (XA-L₀) with a tolerance-region confusion matrix (TOR-Confusion Matrix) to generate minimal input perturbations that induce target predictions while preserving sparsity and semantic interpretability. This unified framework simultaneously enables the generation of highly interpretable counterfactual examples and fine-grained quantification of model robustness. Empirical evaluations on both image and tabular datasets demonstrate the method’s effectiveness, significantly outperforming existing black-box analysis techniques in terms of interpretability and robustness assessment fidelity.
The impact of training data quality on classifier performance is often overlooked. In the context of metagenomic DNA sequence assembly, this study systematically evaluates the behavior of Bayesian classifiers, neural networks, partition models, and random forests under various training data degradation scenarios. The findings reveal that as data quality deteriorates, all classifiers exhibit a “catastrophic” degradation pattern—shifting from substantially correct predictions to essentially random guesses. Concurrently, decision boundaries become sparser, and inter-classifier agreement paradoxically increases, indicating a convergence in error patterns under low-quality training conditions. This work provides the first quantitative characterization of the relationship between data quality and heterogeneity in classifier behavior, offering new insights for designing robust classification systems in data-scarce or noisy environments.