🤖 AI Summary
Traditional cosine similarity assumes data reside in a Euclidean space and ignores the variance and covariance of random variables, leading to inaccurate similarity estimates when features exhibit correlation and heteroscedasticity. To address this, we propose Variance–Covariance-Corrected Cosine distance (VC-Cosine), the first method to explicitly incorporate second-order statistical structure—i.e., feature variances and covariances—into cosine distance computation. VC-Cosine whitens the feature space via the empirical covariance matrix, thereby adaptively reweighting dimensions according to their correlations and variabilities, and aligning the inner-product geometry with the true underlying data distribution rather than relying on isotropic assumptions. Experiments on the Wisconsin Breast Cancer dataset demonstrate that, when integrated with a k-nearest neighbors classifier, VC-Cosine achieves 100% test accuracy—substantially outperforming standard cosine similarity and other state-of-the-art similarity measures.
📝 Abstract
Cosine similarity is a popular distance measure that measures the similarity between two vectors in the inner product space. It is widely used in many data classification algorithms like K-Nearest Neighbors, Clustering etc. This study demonstrates limitations of application of cosine similarity. Particularly, this study demonstrates that traditional cosine similarity metric is valid only in the Euclidean space, whereas the original data resides in a random variable space. When there is variance and correlation in the data, then cosine distance is not a completely accurate measure of similarity. While new similarity and distance metrics have been developed to make up for the limitations of cosine similarity, these metrics are used as substitutes to cosine distance, and do not make modifications to cosine distance to overcome its limitations. Subsequently, we propose a modified cosine similarity metric, where cosine distance is adjusted by variance-covariance of the data. Application of variance-adjusted cosine distance gives better similarity performance compared to traditional cosine distance. KNN modelling on the Wisconsin Breast Cancer Dataset is performed using both traditional and modified cosine similarity measures and compared. The modified formula shows 100% test accuracy on the data.