Score
Designs and implements metrics, tests, and pipelines to quantify and compare the quality, coherence, stability, and interpretability of clustering results. This work includes computing internal and external validation scores, measuring intra‑cluster similarity and inter‑cluster separation, benchmarking clustering variants and resolutions, and producing validated, interpretable segments of records for downstream use.
This paper systematically evaluates the performance of 26 internal clustering validity indices (CVIs) to address the fundamental problem of reliably selecting the optimal clustering solution from a set of candidates. We propose a triple-unbiased methodological framework, wherein each sub-method employs dual complementary metrics—e.g., stability and accuracy—to rigorously assess CVIs along three orthogonal dimensions: robustness, scenario adaptability, and algorithmic independence. Our benchmarking infrastructure comprises 16,177 synthetic and real-world datasets, eight state-of-the-art clustering algorithms, and an enhanced evaluation protocol—constituting the largest CVI benchmark to date. Experimental results reveal systematic strengths and weaknesses of each CVI across diverse data characteristics, including cluster shape, noise level, and dimensionality. The study delivers an interpretable, reproducible, and empirically grounded guideline for CVI selection in practical clustering applications.
Clustering evaluation commonly relies on labeled benchmark datasets, yet their class labels may not accurately reflect the underlying cluster structure, leading to misleading validation. This paper addresses the “cluster–label matching” (CLM) problem by proposing, for the first time, four axiomatic principles to guide the design of invariant, adjusted internal validation measures (Adjusted IVMs). Methodologically, we build upon six classic IVMs—including Silhouette—and introduce a standardized transformation protocol coupled with axiom-driven normalization to ensure cross-dataset comparability and reliability in CLM quantification. Experiments demonstrate that the proposed Adjusted IVMs significantly outperform both original IVMs and state-of-the-art methods in single- and multi-dataset CLM evaluation. Our work provides both theoretical foundations and practical tools for constructing high-fidelity clustering benchmarks.
High-stakes domains—such as healthcare and finance—demand interpretable clustering outcomes to ensure transparency, accountability, and regulatory compliance. Method: This survey systematically analyzes over 120 scholarly works, proposing the first unified taxonomy of interpretability dimensions for clustering. It rigorously distinguishes intrinsically interpretable models—including rule-based, prototype-based, and sparsity-driven approaches—from post-hoc explanation techniques—such as visualization, feature attribution, and local surrogate modeling. The study further develops a use-case-oriented, structured classification framework and principled evaluation criteria. Contribution/Results: It introduces the first practical guideline for selecting appropriate interpretable clustering methods based on application requirements. The work bridges theoretical foundations with real-world deployment, providing both conceptual clarity and actionable insights to support the development and adoption of clustering algorithms that jointly optimize accuracy and interpretability—thereby advancing trustworthy AI in ethically and regulatorily sensitive contexts.
Existing external clustering evaluation metrics (e.g., NMI, ARI) suffer from a lack of monotonicity, inability to identify worst-case scenarios, poor interpretability, and reliance on random baseline assumptions—hindering reliable algorithm comparison. To address these limitations, we propose Normalized Clustering Accuracy (NCA), an asymmetric, normalization-based accuracy measure grounded in optimal set matching. NCA is the first metric to simultaneously satisfy monotonicity, scale invariance, correction for cluster-size imbalance, and sensitivity to worst-case performance. Crucially, it does not rely on adjusted-for-chance assumptions. We provide rigorous theoretical analysis establishing its desirable mathematical properties. Empirical evaluation across multiple benchmark datasets demonstrates that NCA exhibits heightened sensitivity to low-quality clusterings and yields significantly more consistent algorithm rankings than conventional metrics. Thus, NCA provides a more robust, interpretable, and theoretically sound standard for external clustering evaluation when ground-truth labels are available.
This work addresses the critical challenge of effectively comparing structural differences between two entity resolution (ER) clustering results in the absence of ground-truth labels. It proposes the Case Count Metric System (CCMS), which introduces and operationalizes, for the first time, a quantitative framework for four types of cluster transformations—preservation, merging, splitting, and overlapping—without requiring labeled data. By leveraging a cluster-set transformation analysis algorithm, CCMS enables fine-grained, unsupervised comparison of ER outcomes. Integrated with interactive analysis and visualization capabilities, the system has been successfully deployed in both academic and industrial settings, significantly enhancing the interpretability and efficiency of ER method evaluation and tuning.
Clustering analyses often lack quantitative assessment of reproducibility. To address this gap, this work proposes ERICA, the first systematic framework for quantifying clustering reproducibility. ERICA generates stability statistics through iterative cluster assignments and integrates quantitative visualization to reveal inter-cluster similarity and potential outliers. The method is validated on synthetic datasets and applied to breast cancer gene expression data, where it identifies subsets of clustering results that are irreproducible. These findings underscore ERICA’s critical value in real-world applications for evaluating the reliability of clustering outcomes and the robustness of underlying data structures.
This study addresses a critical gap in clustering interpretability: existing post-hoc explanation methods primarily focus on feature importance or instance-level explanations and struggle to reliably uncover structured patterns within clusters. To systematically evaluate this limitation, the authors conduct the first controlled assessment of multiple explanation techniques—including random forest permutation importance, LIME, and principal component analysis—in synthetic datasets where ground-truth structured patterns are explicitly embedded. Results demonstrate that while these methods partially recover relevant features, none consistently identifies all types of predefined patterns. This reveals a fundamental shortcoming of current interpretability tools in capturing pattern-level cluster structure and underscores the urgent need for dedicated methods designed specifically for detecting and explaining such intra-cluster patterns.
In clustering, strong dominance in the size of a particular cluster is often undesirable, motivating a measure of cluster size uniformity that can be used to filter such partitions. A basic requirement of such a measure is stability: partitions that differ only slightly in their point assignments should receive similar uniformity scores. A difficulty arises because cluster labels are not fixed objects; algorithms may produce different numbers of labels even when the underlying point distribution changes very little. Measures defined directly over labels can therefore become unstable under label-count perturbations. I introduce the Mass Agreement Score (MAS), a point-centric metric bounded in [0, 1] that evaluates the consistency of expected cluster size as measured from the perspective of points in each cluster. Its construction yields fragment robustness by design, assigning similar scores to partitions with similar bulk structure while remaining sensitive to genuine redistribution of cluster mass.
Despite the availability of ground-truth labels, systematic guidance on effectively evaluating clustering quality remains limited. This work presents a comprehensive comparison of external clustering evaluation metrics based on set matching, explicitly distinguishing between cluster-level and point-level assessment needs. It advocates for the use of the intuitive and interpretable Centroid Index (CI) for cluster-level evaluation, while highlighting the advantages of metrics such as the Pairwise Set Index (PSI) in point-level contexts. Through an integrated analysis of CI, PSI, clustering accuracy (ACC), and related measures, the study establishes a practical guideline for metric selection tailored to different granularity requirements. This approach substantially enhances the fairness and interpretability of clustering performance evaluation.
Existing clustering evaluation metrics struggle to capture the fine-grained structure inherent in contingency tables and lack a dedicated measure analogous to the confusion matrix in supervised learning. To address this gap, this work proposes the first Associativity-Peakiness (AP) metric specifically designed for contingency tables, which models their intrinsic structural properties to effectively characterize key performance aspects of clustering results. The AP metric exhibits superior dynamic range and computational efficiency compared to existing measures. Experimental evaluation on 500 synthetic contingency tables demonstrates that AP significantly outperforms current metrics in both fine-grained representational capacity and discriminative power, thereby filling a critical void in clustering evaluation methodology.