🤖 AI Summary
This study addresses the lack of image-based contextual information in recommender systems by, for the first time, utilizing images as an independent context source rather than solely for item representation. Methodologically, it leverages vision-language models to extract physical, social, and modal contexts, and proposes the ICE-Fuse framework to integrate these multi-source signals into a review-aware graph contrastive learning algorithm. Empirical results demonstrate that although image context alone underperforms traditional signals, it exhibits significant complementarity with textual information. Fusing these modalities effectively enhances recommendation performance, thereby extending the context modeling paradigm for multimodal recommendation.
📝 Abstract
Contextual information, capturing the circumstances of a user-item interaction, is central to recommender systems. Prior work draws context from location, time, or reviews, but not images; multimodal recommender systems mainly use images to enrich item or user representations, not identify situational context. We propose a new representation of context derived from images, spanning physical, social, and modal categories learned via a vision-language model. We introduce ICE-Fuse, a pipeline for evaluating this representation that fuses these categories and integrates them into a context-aware recommender system, using TripAdvisor data and Review-aware Graph Contrastive Learning as the recommendation algorithm. Image context does not outperform established signals standalone, but improves them combined, indicating complementary information. Semantic analysis shows image- and review-derived context capture distinct aspects of the interaction, positioning images as complementary context.