🤖 AI Summary
This paper addresses the challenge of identifying product substitutability in retail settings. We propose a direct, statistically grounded measure based on co-purchase behavior—specifically, pairwise association metrics (e.g., phi coefficient or lift) derived from co-occurrence frequency matrices—bypassing indirect modeling paradigms rooted in utility theory or embedding-based representations. Our approach ensures computational efficiency, semantic transparency, and strong interpretability. By integrating hierarchical clustering, it enables unsupervised discovery of compact, highly cohesive substitute clusters. Experiments on real-world data from a pharmacy-retail chain demonstrate that our method significantly outperforms baseline approaches and effectively discriminates substitutability from complementarity. The resulting quantifiable, actionable insights support practical applications in category management and shelf-space optimization.
📝 Abstract
We propose a measure of product substitutability based on correlation of common purchases, which is fast to compute and easy to interpret. In an empirical study of a drugstore retail chain, we demonstrate its properties, compare it to a similarly simple measure of product complementarity, and use it to find small clusters of substitutes.