🤖 AI Summary
本文提出了一种非参数框架,使用伪隔离异常分数来定义和检测异常值,并将其应用于K-means聚类中以识别特定于簇的异常值。
📝 Abstract
Outlier detection is a fundamental challenge in data processing, with critical implications for robustness across statistical modeling, machine learning and exploratory data analysis. However, existing proposals rarely offer a universal, domain-agnostic definition of an outlier, often relying on heuristic trimming quotas that lack a statistical interpretation. To address this, we propose a nonparametric framework built on a pseudo-isolation outlier score. This score enables a formal, probabilistic definition of an anomaly tied to a false-alarm rate $\alpha$, which extends into a rigorous geometric classification of internal and external outliers. We show that this mechanism seamlessly embeds into any objective-based clustering framework to identify cluster-specific outliers. Here, we integrate it into $K$-means to create ODK-means. The inferential capabilities and topological properties of this framework are explored both theoretically through formal propositions and empirically through extensive simulations and methodological tutorials, highlighting the practical actionability and intuitive appeal of the proposed outlier detection logic.