Explainable Evidential Clustering

📅 2025-07-16
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Evidence clustering results often lack interpretability in high-stakes domains (e.g., healthcare), hindering trust and adoption. Method: This paper proposes an explainable modeling framework grounded in Dempster–Shafer theory. Unlike conventional approaches, it employs representative conditions as primitives and integrates utility functions with evidence misclassification costs to construct an interpretable decision tree tailored for evidence classifiers. Furthermore, it introduces the Iterative Evidence Misclassification Minimization (IEMM) algorithm to generate cautious, transparent, and preference-aligned cluster explanations. Contribution/Results: Experiments on synthetic and real-world datasets demonstrate that the method substantially improves explanation quality, achieving a 93% explanation satisfaction rate. It establishes a novel paradigm for trustworthy clustering decisions under uncertainty, advancing both interpretability and decision-theoretic rigor in evidence-based analytics.

Technology Category

Data Mining & Knowledge Management: Representing, Reasoning, and Using Provenance, TrustReasoning under Uncertainty: Other Foundations of Reasoning under UncertaintyMachine Learning: Clustering

Application Category

User Modeling, Personalization and Recommendation: Explainable and interpretable methods for personalizationWeb Mining and Content Analysis: Normalization, clustering, classification, and summarization of Web textSemantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactions
📝 Abstract
Unsupervised classification is a fundamental machine learning problem. Real-world data often contain imperfections, characterized by uncertainty and imprecision, which are not well handled by traditional methods. Evidential clustering, based on Dempster-Shafer theory, addresses these challenges. This paper explores the underexplored problem of explaining evidential clustering results, which is crucial for high-stakes domains such as healthcare. Our analysis shows that, in the general case, representativity is a necessary and sufficient condition for decision trees to serve as abductive explainers. Building on the concept of representativity, we generalize this idea to accommodate partial labeling through utility functions. These functions enable the representation of "tolerable" mistakes, leading to the definition of evidential mistakeness as explanation cost and the construction of explainers tailored to evidential classifiers. Finally, we propose the Iterative Evidential Mistake Minimization (IEMM) algorithm, which provides interpretable and cautious decision tree explanations for evidential clustering functions. We validate the proposed algorithm on synthetic and real-world data. Taking into account the decision-maker's preferences, we were able to provide an explanation that was satisfactory up to 93% of the time.
Problem

Research questions and friction points this paper is trying to address.

Explaining evidential clustering results for high-stakes domains
Addressing imperfections like uncertainty in unsupervised classification
Developing interpretable decision tree explanations for evidential classifiers
Innovation

Methods, ideas, or system contributions that make the work stand out.

Evidential clustering based on Dempster-Shafer theory
Decision trees as abductive explainers via representativity
IEMM algorithm for interpretable cautious explanations
🔎 Similar Papers
2024-09-01arXiv.orgCitations: 4