Interpretable Clustering Ensemble

📅 2025-06-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
High-stakes domains such as medical diagnosis and financial risk management demand interpretable clustering ensembles, yet existing methods lack transparency in decision logic. Method: This paper proposes the first interpretable clustering ensemble framework, which models base clustering results as categorical variables and constructs decision trees directly in the original feature space—guided by statistical association tests (e.g., chi-square test)—to explicitly encode and trace clustering rationale. A co-designed categorical encoding scheme and ensemble strategy jointly ensure both fidelity and interpretability. Contribution/Results: The method achieves state-of-the-art clustering quality across multiple benchmark datasets while offering human-readable, auditable clustering rules. It bridges a critical gap in the literature by establishing the first principled approach to interpretable clustering ensembles, advancing both theoretical understanding and practical deployment in safety-critical applications.

Technology Category

Machine Learning: ClusteringComputer Vision: Interpretability, Explainability, and TransparencyKnowledge Representation and Reasoning: Diagnosis and Abductive Reasoning

Application Category

Web Mining and Content Analysis: Normalization, clustering, classification, and summarization of Web textUser Modeling, Personalization and Recommendation: Explainable and interpretable methods for personalizationSemantics and Knowledge: Methods, algorithms and applications for the development of semantic models, knowledge graphs and other forms of structured data models with machine-interpretable semantics
📝 Abstract
Clustering ensemble has emerged as an important research topic in the field of machine learning. Although numerous methods have been proposed to improve clustering quality, most existing approaches overlook the need for interpretability in high-stakes applications. In domains such as medical diagnosis and financial risk assessment, algorithms must not only be accurate but also interpretable to ensure transparent and trustworthy decision-making. Therefore, to fill the gap of lack of interpretable algorithms in the field of clustering ensemble, we propose the first interpretable clustering ensemble algorithm in the literature. By treating base partitions as categorical variables, our method constructs a decision tree in the original feature space and use the statistical association test to guide the tree building process. Experimental results demonstrate that our algorithm achieves comparable performance to state-of-the-art (SOTA) clustering ensemble methods while maintaining an additional feature of interpretability. To the best of our knowledge, this is the first interpretable algorithm specifically designed for clustering ensemble, offering a new perspective for future research in interpretable clustering.
Problem

Research questions and friction points this paper is trying to address.

Lack of interpretable algorithms in clustering ensemble
Need for transparent decision-making in high-stakes applications
Combining clustering accuracy with interpretability in ensemble methods
Innovation

Methods, ideas, or system contributions that make the work stand out.

First interpretable clustering ensemble algorithm
Decision tree in original feature space
Statistical association test guides tree building
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
H
Hang Lv
School of Software, Dalian University of Technology, Dalian, China
L
Lianyu Hu
School of Software, Dalian University of Technology, Dalian, China
M
Mudi Jiang
School of Software, Dalian University of Technology, Dalian, China
X
Xinying Liu
School of Software, Dalian University of Technology, Dalian, China
Z
Zengyou He
School of Software, Dalian University of Technology, Dalian, China