Score
Designs and evaluates methods that learn dissimilarity metrics for categorical attributes, producing intra-attribute distance functions and combined measures that represent both nominal categories and ordinal ordering. This includes building graph-based intra-attribute distance representations, order-preserving distance transforms, and metrics that model nominal–ordinal interdependence for use in downstream analysis or learning.
Existing clustering methods fail to distinguish the semantic differences between nominal and ordinal attributes, often neglecting the inherent ordering among ordinal values and inter-attribute dependencies, which leads to inaccurate distance measurements. To address this limitation, this work proposes a unified graph-based distance metric framework that explicitly models the intrinsic structures of both attribute types. The approach introduces learnable intra-attribute distance weights and jointly optimizes distance learning with cluster assignment in an end-to-end manner. Notably, it is the first method to simultaneously preserve ordinal information and adaptively adjust attribute weights within a single paradigm. Extensive experiments on multiple real-world datasets demonstrate that the proposed method significantly outperforms state-of-the-art algorithms, confirming its effectiveness and superiority.
Categorical data lack a natural metric space, posing a fundamental challenge for distribution modeling and clustering due to the absence of meaningful distance measures. To address this, we propose a novel distance metric learning paradigm centered on the ordinal relationships among attribute values. We theoretically establish that value ordering—not raw categorical similarity—is the essential determinant of clustering accuracy, and we formalize the link between ordinal structure and cluster membership. Our method jointly optimizes value ordering and clustering assignments via an alternating optimization framework, yielding a provably convergent, interpretable distance function. Extensive experiments on multiple benchmark datasets demonstrate significant improvements over state-of-the-art methods. Ablation studies and statistical tests confirm that learned ordinal relationships consistently enhance both clustering accuracy and semantic interpretability of resulting clusters.
Existing categorical clustering methods suffer from two key limitations: (i) the absence of prior semantic relationships among categories, and (ii) the rigid assumption of fixed-distance metrics, which poorly accommodates diverse cluster structures. To address these, this paper proposes a cluster-adaptive, learnable distance metric. Our core contributions are threefold: (i) the first formulation of *cluster-specific* categorical relationship modeling—eliminating reliance on predefined topological structures; (ii) a differentiable distance framework that jointly optimizes categorical relationships and clustering objectives via end-to-end learning; and (iii) native compatibility with Euclidean distance, enabling seamless extension to mixed-type data. Extensive experiments across 12 real-world categorical datasets demonstrate state-of-the-art performance: our method achieves a mean clustering accuracy rank of 1.25, substantially outperforming the current best approach (rank 5.21), thereby validating its superior capacity to model complex distributions and generalize across heterogeneous data.
This paper addresses the performance degradation of conventional decision trees in ordinal classification (OC) tasks, stemming from their neglect of label-order relationships. To remedy this, we systematically design and evaluate order-aware splitting criteria. We introduce a unified notation and conduct the first large-scale empirical comparison—across a benchmark comprising 45 publicly available OC datasets—of several ordinal-specific criteria, including Ordinal Gini (OGini), Weighted Information Gain, and Ranking Impurity. Results demonstrate that OGini significantly outperforms nominal Gini and standard information gain, establishing it as the current state-of-the-art ordinal splitting criterion. Our approach integrates seamlessly into standard decision tree frameworks and employs ordinal-sensitive evaluation metrics such as MAE and ORMSE. All code, datasets, and experimental results are fully open-sourced to foster reproducible research in ordinal learning.
This work addresses the limitation of the Tanimoto kernel (Jaccard index), which is restricted to binary sets or nonnegative real-valued functions. We propose the first generalized Tanimoto kernel for **arbitrary real-valued functions**. Our method decomposes each function into signed magnitude components, mapping it to a pair of signed sets; this yields a rigorously defined set-based representation, from which we derive an explicit feature map and the associated reproducing kernel Hilbert space (RKHS) structure. Building on general kernel design principles, we further provide a piecewise-linear analytic formulation and a differentiable smooth approximation. The resulting framework unifies similarity modeling for real-valued functions, combining theoretical soundness with computational tractability. Empirically, it significantly improves generalization performance in function regression and similarity learning tasks.
This work addresses the performance degradation in ordinal classification caused by existing methods' neglect of the natural order among classes. To this end, we propose ADABORD, a novel framework that, for the first time, integrates both an ordinal splitting criterion and an error function accounting for inter-class distances within AdaBoost. Specifically, ADABORD employs decision stumps based on an ordinal Gini impurity measure as base learners and introduces an absolute ranking probability score to more appropriately update sample and model weights. Experimental results on the TOC-UCO benchmark—the largest evaluation suite for ordinal classification—demonstrate that ADABORD significantly outperforms seven state-of-the-art methods, with particularly pronounced gains on datasets containing five or more ordinal classes.
This work addresses the limitation of traditional clustering methods that rely on Euclidean distance and struggle to uncover implicit cluster structures in qualitative data, such as symptoms or marital status. To overcome this, the authors propose an ordered forest representation tailored for clustering tasks, which— for the first time—models nominal attribute values as vertices in tree structures to capture their flexible local ordinal relationships. A clustering-oriented joint learning framework is further designed to simultaneously optimize both the tree structures and cluster assignments. Extensive experiments on twelve real-world benchmark datasets demonstrate that the proposed method significantly outperforms ten state-of-the-art baselines, confirming its effectiveness and superiority in clustering qualitative data.
This study addresses the lack of systematic and neutral comparisons among similarity measures for categorical datasets. It presents the first comprehensive evaluation of several prominent methods—including edge-count tests, constrained minimum distance, graph-based tests, Classifier Two-Sample Tests (C2ST), and the Maximum Mean Discrepancy with Categorical Metrics (MMCM)—assessing their ability to detect distributional differences and their computational costs in both two-sample and multi-sample settings. The results demonstrate that the Friedman–Rafsky test achieves the best overall performance in two-sample tasks, while MMCM excels in multi-sample scenarios by offering both high statistical power and computational efficiency. This work provides empirical evidence and practical guidance for selecting appropriate similarity measures when analyzing categorical data.
This work addresses the lack of general-purpose methods and open-source tools for ordinal classification by proposing a model-agnostic framework that transforms any base classifier into an ordinal-aware variant. The approach integrates a classifier pooling strategy with ordinal constraint mechanisms, enabling, for the first time, universal adaptation of arbitrary classifiers to ordinal data. To support reproducibility and practical adoption, the authors release an open-source Python package that fills a critical gap in available ordinal classification tooling. Extensive experiments on multiple real-world datasets demonstrate that the proposed method significantly outperforms conventional non-ordinal classifiers, particularly in small-sample and high-cardinality settings, thereby confirming its effectiveness and practical utility.