OpticalRec: Unified Optical Vision-Language Representation for Multimodal Recommendation

📅 2026-10-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the absence of visual-linguistic semantic interaction and cross-modal distortion caused by independent encoding in multimodal recommendation. To overcome these limitations, we propose the first unified visual-space encoding paradigm. By rendering text as visual glyphs and employing a dual-attention mechanism, this approach enables native image-text interaction at both perceptual and semantic levels, thereby transcending traditional late-fusion bottlenecks. The architecture comprises a visual encoder, a language decoder, and plug-and-play modules, supported theoretically through mutual information analysis. Experimental results demonstrate that the proposed method significantly outperforms strong baselines while exhibiting low computational cost, high robustness, and seamless integration capabilities.
📝 Abstract
Recent advances in vision-language modeling have substantially improved multimodal encoding, retrieval and reasoning. Yet for multimodal recommendation, encoding rich item vision-language semantic interactions remains a long-standing bottleneck, which hampers accurate item representation learning and user-item matching. Mainstream approaches primarily adopt independent encoding of vision and language modality followed by rigid late fusion such as concatenation, inherently omitting native vision-language interactions and introducing cross-modal semantic distortion. To address this challenge, we propose OpticalRec, the first visual-space unified encoding paradigm for multimodal collaborative filtering, a fundamental recommendation setting. Instead of isolated modality-specific encoding, OpticalRec renders item textual metadata as visual glyphs, enabling native image-text interaction within the visual encoder - the perceptual encoding level. The resulting representations are further processed by the language decoder - the semantic encoding level, allowing OpticalRec to exploit the dual-attention mechanism of modern vision-language models that previous encoding methods omitted. OpticalRec's efficacy is theoretically supported by mutual information analysis and empirically demonstrated through superior performance across strong baselines and benchmarks. As a plug-and-play module, OpticalRec (1) introduces minimal cost, (2) is robust against rendered text font, color and layout, etc., and (3) integrates seamlessly into existing multimodal collaborative filtering models.
Problem

Research questions and friction points this paper is trying to address.

multimodal recommendation
vision-language interaction
cross-modal semantic distortion
item representation learning
collaborative filtering
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multimodal Recommendation
Unified Encoding
Vision-Language Model
Collaborative Filtering
Visual Glyphs
🔎 Similar Papers
2024-08-08International Workshop on Semantic and Social Media Adaptation and PersonalizationCitations: 13