Variance-Adjusted Cosine Distance as Similarity Metric

📅 2025-02-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Traditional cosine similarity assumes data reside in a Euclidean space and ignores the variance and covariance of random variables, leading to inaccurate similarity estimates when features exhibit correlation and heteroscedasticity. To address this, we propose Variance–Covariance-Corrected Cosine distance (VC-Cosine), the first method to explicitly incorporate second-order statistical structure—i.e., feature variances and covariances—into cosine distance computation. VC-Cosine whitens the feature space via the empirical covariance matrix, thereby adaptively reweighting dimensions according to their correlations and variabilities, and aligning the inner-product geometry with the true underlying data distribution rather than relying on isotropic assumptions. Experiments on the Wisconsin Breast Cancer dataset demonstrate that, when integrated with a k-nearest neighbors classifier, VC-Cosine achieves 100% test accuracy—substantially outperforming standard cosine similarity and other state-of-the-art similarity measures.

Technology Category

Machine Learning: Dimensionality Reduction/Feature SelectionComputer Vision: Representation Learning for VisionCognitive Modeling & Cognitive Systems: Analogy

Application Category

Web Mining and Content Analysis: Robustness and generalizability of Web mining methodsUser Modeling, Personalization and Recommendation: Fairness-aware retrieval and rankingGraph Algorithms and Modeling for the Web: Graph embeddings and representation learning for Web-related graphs
📝 Abstract
Cosine similarity is a popular distance measure that measures the similarity between two vectors in the inner product space. It is widely used in many data classification algorithms like K-Nearest Neighbors, Clustering etc. This study demonstrates limitations of application of cosine similarity. Particularly, this study demonstrates that traditional cosine similarity metric is valid only in the Euclidean space, whereas the original data resides in a random variable space. When there is variance and correlation in the data, then cosine distance is not a completely accurate measure of similarity. While new similarity and distance metrics have been developed to make up for the limitations of cosine similarity, these metrics are used as substitutes to cosine distance, and do not make modifications to cosine distance to overcome its limitations. Subsequently, we propose a modified cosine similarity metric, where cosine distance is adjusted by variance-covariance of the data. Application of variance-adjusted cosine distance gives better similarity performance compared to traditional cosine distance. KNN modelling on the Wisconsin Breast Cancer Dataset is performed using both traditional and modified cosine similarity measures and compared. The modified formula shows 100% test accuracy on the data.
Problem

Research questions and friction points this paper is trying to address.

Limitations of traditional cosine similarity in random variable space
Impact of variance and correlation on cosine distance accuracy
Proposal of a variance-adjusted cosine similarity for improved performance
Innovation

Methods, ideas, or system contributions that make the work stand out.

Variance-adjusted cosine similarity
Improved KNN classification accuracy
Data variance-covariance integration
💼 Related Jobs
No related jobs found.
IIT Kharagpur