Seeing the Context: Enhancing Recommender Systems with Image-Derived Contextual Signals

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of image-based contextual information in recommender systems by, for the first time, utilizing images as an independent context source rather than solely for item representation. Methodologically, it leverages vision-language models to extract physical, social, and modal contexts, and proposes the ICE-Fuse framework to integrate these multi-source signals into a review-aware graph contrastive learning algorithm. Empirical results demonstrate that although image context alone underperforms traditional signals, it exhibits significant complementarity with textual information. Fusing these modalities effectively enhances recommendation performance, thereby extending the context modeling paradigm for multimodal recommendation.
📝 Abstract
Contextual information, capturing the circumstances of a user-item interaction, is central to recommender systems. Prior work draws context from location, time, or reviews, but not images; multimodal recommender systems mainly use images to enrich item or user representations, not identify situational context. We propose a new representation of context derived from images, spanning physical, social, and modal categories learned via a vision-language model. We introduce ICE-Fuse, a pipeline for evaluating this representation that fuses these categories and integrates them into a context-aware recommender system, using TripAdvisor data and Review-aware Graph Contrastive Learning as the recommendation algorithm. Image context does not outperform established signals standalone, but improves them combined, indicating complementary information. Semantic analysis shows image- and review-derived context capture distinct aspects of the interaction, positioning images as complementary context.
Problem

Research questions and friction points this paper is trying to address.

Recommender Systems
Contextual Information
Image Context
Multimodal Recommendation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Context-aware Recommender Systems
Vision-Language Model
Image-derived Context
Multimodal Recommendation
Graph Contrastive Learning
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
T
Tal Cordova
Coller School of Management, Tel Aviv University, Israel
T
Tomer Geva
Coller School of Management, Tel Aviv University, Israel
Moshe Unger
Moshe Unger
Assistant Professor at Tel Aviv University - Coller School of Management
Data ScienceRecommender SystemsMachine LearningCyber SecurityDeep Learning