Score
Designs and implements metrics and analysis methods that estimate a dataset's difficulty for supervised tasks—particularly classification—by producing data-side difficulty scores, single-number summaries for rapid dataset comparison, measures of few-class distinctiveness, and predictors of expected model accuracy without full training. This work quantifies dataset properties such as class separability, noise, sample complexity, and feature overlap to compare datasets and anticipate model performance with minimal or no training.
Existing dataset characterization methods—statistical, structural, and model-driven—lack sufficient interpretability and deep structural insight. To address this, we propose a novel tensor-based representation paradigm that transcends conventional two-dimensional assumptions. Our approach leverages high-order tensor decomposition, multilinear modeling, and cross-modal joint representation to explicitly capture high-dimensional, nonlinear, and multi-source relational structures inherent in complex data. Extensive experiments demonstrate that the proposed method significantly outperforms baseline approaches in three key aspects: (i) disentangling intricate data structures, (ii) enhancing feature interpretability, and (iii) enabling traceable downstream task reasoning. This work establishes a unified tensor modeling framework for dataset representation and pioneers a data-driven discovery pathway tailored for explainable AI. It contributes both theoretical advances—through formalizing multilinear structure learning—and practical utility—by providing an interpretable, computationally grounded toolkit for transparent data analysis.
Current evaluation practices for supervised learning models are often misleading due to an overreliance on single aggregate metrics, which neglect the alignment among data characteristics, task objectives, and real-world application contexts. This work reframes model evaluation as a context-dependent, decision-oriented process and systematically investigates—through controlled experiments—the impact of dataset properties, validation strategies, class imbalance, and asymmetric error costs on evaluation outcomes. Leveraging diverse benchmark datasets, multiple validation protocols, and multidimensional performance measures, the study uncovers common pitfalls such as the accuracy paradox, data leakage, and metric misuse. It proposes a structured evaluation framework explicitly aligned with operational goals, offering principled guidance for developing more robust, reliable, and trustworthy supervised learning systems.
The few-shot and imbalanced (S&I) learning problem suffers from severe generalization degradation and low interpretability due to scarce samples, extreme class imbalance, and ambiguous inter-class feature distributions. This paper proposes the first systematic analytical framework tailored to S&I learning, advocating that quantitative characterization of data properties—such as imbalance ratio and geometric complexity—must precede algorithmic design. The framework unifies multi-dimensional imbalance metrics, data complexity analysis, resampling strategies, classifier adaptation mechanisms, and an interpretable evaluation benchmark. Empirical evaluation on binary and multi-class extreme imbalance benchmarks reveals that classifier selection exerts significantly greater impact on performance than resampling improvements—exposing a fundamental flaw in prevailing heuristic-driven approaches. Our work establishes a theory-guided analytical paradigm and practical design principles for S&I learning, advancing both methodological rigor and empirical reproducibility.
This study addresses the critical need for quantifying dataset similarity in model generalization, transfer learning, simulation calibration, and two-sample testing. We systematically survey 118 similarity quantification methods and propose the first ten-dimensional classification framework, organizing approaches into seven technical categories: statistical distances (e.g., Wasserstein distance, Maximum Mean Discrepancy), kernel-based methods, information-theoretic measures, dimensionality-reduction embeddings, permutation tests, generative-model-based discriminators, and Gaussian process likelihood ratios. We develop a multi-dimensional evaluation system balancing theoretical guarantees, interpretability, and practical applicability, yielding a structured recommendation matrix aligned with task requirements and data characteristics. Furthermore, we introduce the first open-source, interactive tool for method selection—enabling real-time filtering and parameter configuration—to significantly enhance both selection efficiency and deployment suitability.
This paper addresses the challenge of quantifying the intrinsic difficulty of a dataset for a given model. Method: We propose an information-theoretic framework—V-usable information—and introduce Pointwise V-Information (PVI), a fine-grained, model-aware metric that formalizes sample-level difficulty as the deficit of information usable by the model. PVI enables interpretable difficulty attribution across datasets, subpopulations, and input attributes, supports reverse evaluation (fixed model, multiple datasets), and facilitates diagnosis of annotation artifacts. Contribution/Results: Evaluated on NLP benchmarks, the framework successfully uncovers latent annotation biases. Experiments confirm its model-agnosticism, cross-task consistency, and strong interpretability. By grounding data difficulty in information usability, PVI provides a unified, computationally tractable theoretical tool for rigorous difficulty analysis.
This study investigates whether effect size measures (e.g., Cohen’s *d*) can serve as prospective proxies for data sufficiency—specifically, to predict model performance (classification accuracy) and training convergence speed. We conduct systematic supervised learning experiments across varying sample sizes, learning rates, and convergence dynamics, quantitatively assessing statistical associations between effect size and model behavior. Our first empirical evaluation reveals no robust correlation between effect size and either accuracy or convergence rate, demonstrating its unreliability for sample-size planning or performance forecasting. These findings expose fundamental limitations of conventional descriptive statistics in assessing data sufficiency, challenge the implicit assumption that effect size serves as a valid proxy for data quality, and underscore the need for a new evaluation framework integrating statistical learning theory with explicit modeling of data-generating mechanisms.
This study addresses the inefficiency of neural network model selection in real-world scenarios involving few classes (<10). To tackle this challenge, the authors propose a data-centric metric for quantifying classification difficulty, introducing the novel concept of “few-class discriminability.” This enables the construction of an efficient evaluation framework that facilitates cross-model and cross-dataset comparisons without repeated training. Integrated with lightweight network scaling and cross-platform deployment strategies, the proposed method achieves rapid model selection in applications such as mobile robotics, drones, and IoT systems, accelerating the selection process by 6–29×. Furthermore, it yields a highly compact model that is 42% smaller than YOLOv5-nano while maintaining competitive accuracy and substantially reducing computational resource requirements.
This study addresses the lack of systematic and impartial evaluation of existing methods for measuring distributional similarity in numerical data, which hinders informed selection in practice. The authors construct the first comprehensive benchmarking framework encompassing 36 similarity measures for continuous data—including statistical tests, distance-based metrics, and embedding approaches—and evaluate their discriminative power and computational efficiency through large-scale simulations across diverse distributional discrepancies (e.g., shifts in location, scale, and higher-order moments) and both two-sample and multi-sample settings. Based on empirical performance, the work proposes a data-characteristic-driven strategy for method selection, establishes a performance ranking, and demonstrates that combining only four to six methods suffices to achieve near-optimal performance in 90%–95% of scenarios.
This study addresses the lack of a unified definition and quantification of training dataset diversity, which hinders accurate assessment of its impact on model robustness. Focusing on medical imaging, the work presents the first systematic comparison between reference-free and semantic diversity metrics. Leveraging the MorphoMNIST and PadChest datasets, the authors conduct a multidimensional analysis—integrating Fréchet Inception Distance (FID), AUC, semantic diversity measures, controlled perturbations, and clinical expert evaluations—to examine how diversity correlates with expert intuition, downstream performance, and training dynamics. The findings reveal that FID and semantic diversity more effectively predict model performance, whereas merely increasing the number of imaging device sources can inadvertently encourage models to rely on non-robust shortcut features, thereby exposing a critical pitfall in data diversity design.
This work addresses the challenge of selecting high-quality data subsets from noisy labels to achieve performance approaching that of noise-free training. The authors observe that conventional k-nearest neighbors (k-NN) suffer degraded performance in high-dimensional, label-noisy settings and propose a novel approach that integrates symmetry and invariance priors into subset selection. Specifically, they introduce symmetry into the cutstats framework for the first time and theoretically demonstrate that leveraging invariance enables k-NN to asymptotically approach the Bayes optimal classifier. Moreover, they show that even with only partial knowledge of symmetries, effective modeling is achievable through learned symmetry-aware representations. Empirical results confirm that the proposed method substantially improves subset selection quality under high-dimensional label noise, yielding downstream model performance close to that attainable with clean labels.
This work proposes a unified measure of model complexity based on gradient similarity under input perturbations, applicable to both parametric and non-parametric models. Existing complexity metrics often struggle to balance theoretical rigor with computational efficiency; in contrast, the proposed measure achieves both while offering a coherent framework that subsumes classical notions such as polynomial degree and kernel lengthscale. The approach provides a novel interpretation of the double descent phenomenon by revealing how model complexity evolves throughout the training process. Through rigorous theoretical analysis and empirical validation across diverse model classes—including neural networks and random Fourier features—the method demonstrates strong mathematical grounding and practical scalability, effectively capturing the nuanced dynamics of complexity in modern machine learning systems.