categorical metric learning

Designs and evaluates methods that learn dissimilarity metrics for categorical attributes, producing intra-attribute distance functions and combined measures that represent both nominal categories and ordinal ordering. This includes building graph-based intra-attribute distance representations, order-preserving distance transforms, and metrics that model nominal–ordinal interdependence for use in downstream analysis or learning.

categoricalmetriclearning

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.05
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Existing clustering methods fail to distinguish the semantic differences between nominal and ordinal attributes, often neglecting the inherent ordering among ordinal values and inter-attribute dependencies, which leads to inaccurate distance measurements. To address this limitation, this work proposes a unified graph-based distance metric framework that explicitly models the intrinsic structures of both attribute types. The approach introduces learnable intra-attribute distance weights and jointly optimizes distance learning with cluster assignment in an end-to-end manner. Notably, it is the first method to simultaneously preserve ordinal information and adaptively adjust attribute weights within a single paradigm. Extensive experiments on multiple real-world datasets demonstrate that the proposed method significantly outperforms state-of-the-art algorithms, confirming its effectiveness and superiority.

categorical data clusteringdistance metricintra-attribute distances

Categorical Data Clustering via Value Order Estimated Distance Metric Learning

Nov 19, 2024
YZ
Yiqun Zhang
🏛️ Guangdong University of Technology | Shenzhen University | Hong Kong Baptist University

Categorical data lack a natural metric space, posing a fundamental challenge for distribution modeling and clustering due to the absence of meaningful distance measures. To address this, we propose a novel distance metric learning paradigm centered on the ordinal relationships among attribute values. We theoretically establish that value ordering—not raw categorical similarity—is the essential determinant of clustering accuracy, and we formalize the link between ordinal structure and cluster membership. Our method jointly optimizes value ordering and clustering assignments via an alternating optimization framework, yielding a provably convergent, interpretable distance function. Extensive experiments on multiple benchmark datasets demonstrate significant improvements over state-of-the-art methods. Ablation studies and statistical tests confirm that learned ordinal relationships consistently enhance both clustering accuracy and semantic interpretability of resulting clusters.

Clustering categorical data lacks natural distance metrics.Joint learning of clusters and value orders enhances understanding.Order relation among attribute values improves clustering accuracy.

Break the Tie: Learning Cluster-Customized Category Relationships for Categorical Data Clustering

Nov 12, 2025
MZ
Mingjie Zhao
🏛️ Hong Kong Baptist University | Guangdong University of Technology | Xiamen University | Shenzhen University | BNU-HKBU United International College

Existing categorical clustering methods suffer from two key limitations: (i) the absence of prior semantic relationships among categories, and (ii) the rigid assumption of fixed-distance metrics, which poorly accommodates diverse cluster structures. To address these, this paper proposes a cluster-adaptive, learnable distance metric. Our core contributions are threefold: (i) the first formulation of *cluster-specific* categorical relationship modeling—eliminating reliance on predefined topological structures; (ii) a differentiable distance framework that jointly optimizes categorical relationships and clustering objectives via end-to-end learning; and (iii) native compatibility with Euclidean distance, enabling seamless extension to mixed-type data. Extensive experiments across 12 real-world categorical datasets demonstrate state-of-the-art performance: our method achieves a mean clustering accuracy rank of 1.25, substantially outperforming the current best approach (rank 5.21), thereby validating its superior capacity to model complex distributions and generalize across heterogeneous data.

Breaking fixed category relationships to enhance clustering adaptabilityEnabling seamless extension to mixed numerical-categorical datasetsLearning customized distance metrics for categorical attributes

Splitting criteria for ordinal decision trees: an experimental study

Dec 18, 2024
RA
Rafael Ayll'on-Gavil'an
🏛️ IMIBIC | Universidad Loyola Andalucia | University of Cordoba

This paper addresses the performance degradation of conventional decision trees in ordinal classification (OC) tasks, stemming from their neglect of label-order relationships. To remedy this, we systematically design and evaluate order-aware splitting criteria. We introduce a unified notation and conduct the first large-scale empirical comparison—across a benchmark comprising 45 publicly available OC datasets—of several ordinal-specific criteria, including Ordinal Gini (OGini), Weighted Information Gain, and Ranking Impurity. Results demonstrate that OGini significantly outperforms nominal Gini and standard information gain, establishing it as the current state-of-the-art ordinal splitting criterion. Our approach integrates seamlessly into standard decision tree frameworks and employs ordinal-sensitive evaluation metrics such as MAE and ORMSE. All code, datasets, and experimental results are fully open-sourced to foster reproducible research in ordinal learning.

Compares ordinal and nominal splitting methodsDevelops ordinal decision tree criteriaIdentifies most effective ordinal splitting criterion

On the generalization of Tanimoto-type kernels to real valued functions

Jul 12, 2020
SS
Sandor Szedmak
🏛️ Aalto University

This work addresses the limitation of the Tanimoto kernel (Jaccard index), which is restricted to binary sets or nonnegative real-valued functions. We propose the first generalized Tanimoto kernel for **arbitrary real-valued functions**. Our method decomposes each function into signed magnitude components, mapping it to a pair of signed sets; this yields a rigorously defined set-based representation, from which we derive an explicit feature map and the associated reproducing kernel Hilbert space (RKHS) structure. Building on general kernel design principles, we further provide a piecewise-linear analytic formulation and a differentiable smooth approximation. The resulting framework unifies similarity modeling for real-valued functions, combining theoretical soundness with computational tractability. Empirically, it significantly improves generalization performance in function regression and similarity learning tasks.

Extends Tanimoto kernel to arbitrary real-valued functionsProvides explicit feature representation and smooth approximationUnifies attribute representation via properly chosen sets

Latest Papers

What's happening recently
View more

This work addresses the performance degradation in ordinal classification caused by existing methods' neglect of the natural order among classes. To this end, we propose ADABORD, a novel framework that, for the first time, integrates both an ordinal splitting criterion and an error function accounting for inter-class distances within AdaBoost. Specifically, ADABORD employs decision stumps based on an ordinal Gini impurity measure as base learners and introduces an absolute ranking probability score to more appropriately update sample and model weights. Experimental results on the TOC-UCO benchmark—the largest evaluation suite for ordinal classification—demonstrate that ADABORD significantly outperforms seven state-of-the-art methods, with particularly pronounced gains on datasets containing five or more ordinal classes.

AdaBoostClass OrderNominal Classification

This work addresses the limitation of traditional clustering methods that rely on Euclidean distance and struggle to uncover implicit cluster structures in qualitative data, such as symptoms or marital status. To overcome this, the authors propose an ordered forest representation tailored for clustering tasks, which— for the first time—models nominal attribute values as vertices in tree structures to capture their flexible local ordinal relationships. A clustering-oriented joint learning framework is further designed to simultaneously optimize both the tree structures and cluster assignments. Extensive experiments on twelve real-world benchmark datasets demonstrate that the proposed method significantly outperforms ten state-of-the-art baselines, confirming its effectiveness and superiority in clustering qualitative data.

cluster distributionclusteringdistance structure

This study addresses the lack of systematic and neutral comparisons among similarity measures for categorical datasets. It presents the first comprehensive evaluation of several prominent methods—including edge-count tests, constrained minimum distance, graph-based tests, Classifier Two-Sample Tests (C2ST), and the Maximum Mean Discrepancy with Categorical Metrics (MMCM)—assessing their ability to detect distributional differences and their computational costs in both two-sample and multi-sample settings. The results demonstrate that the Friedman–Rafsky test achieves the best overall performance in two-sample tasks, while MMCM excels in multi-sample scenarios by offering both high statistical power and computational efficiency. This work provides empirical evidence and practical guidance for selecting appropriate similarity measures when analyzing categorical data.

categorical datadataset similaritymulti-sample comparison

This work addresses the lack of general-purpose methods and open-source tools for ordinal classification by proposing a model-agnostic framework that transforms any base classifier into an ordinal-aware variant. The approach integrates a classifier pooling strategy with ordinal constraint mechanisms, enabling, for the first time, universal adaptation of arbitrary classifiers to ordinal data. To support reproducibility and practical adoption, the authors release an open-source Python package that fills a critical gap in available ordinal classification tooling. Extensive experiments on multiple real-world datasets demonstrate that the proposed method significantly outperforms conventional non-ordinal classifiers, particularly in small-sample and high-cardinality settings, thereby confirming its effectiveness and practical utility.

classifier poolingmachine learningmodel-agnostic

Hot Scholars

DZ

Donghuo Zeng

KDDI Research,Inc.
Audio-visual learningCausality
TB

Thomas Bäck

Professor of Computer Science, Leiden University; Chief Scientist, NORCE Research Centre, Norway
Evolutionary ComputationEvolutionary AlgorithmsMachine LearningIndustry 4.0
MD

Ming Dai

SouthEast University
MLLMVisual GroundingImage Retrieval
YL

Yang Lu

Chair Professor of Nanomechanics, Department of Mechanical Engineering, the University of Hong Kong
NanomechanicsNanomanufacturingMechanical MetamaterialsDiamond