🤖 AI Summary
This study addresses the absence of visual-linguistic semantic interaction and cross-modal distortion caused by independent encoding in multimodal recommendation. To overcome these limitations, we propose the first unified visual-space encoding paradigm. By rendering text as visual glyphs and employing a dual-attention mechanism, this approach enables native image-text interaction at both perceptual and semantic levels, thereby transcending traditional late-fusion bottlenecks. The architecture comprises a visual encoder, a language decoder, and plug-and-play modules, supported theoretically through mutual information analysis. Experimental results demonstrate that the proposed method significantly outperforms strong baselines while exhibiting low computational cost, high robustness, and seamless integration capabilities.
📝 Abstract
Recent advances in vision-language modeling have substantially improved multimodal encoding, retrieval and reasoning. Yet for multimodal recommendation, encoding rich item vision-language semantic interactions remains a long-standing bottleneck, which hampers accurate item representation learning and user-item matching. Mainstream approaches primarily adopt independent encoding of vision and language modality followed by rigid late fusion such as concatenation, inherently omitting native vision-language interactions and introducing cross-modal semantic distortion. To address this challenge, we propose OpticalRec, the first visual-space unified encoding paradigm for multimodal collaborative filtering, a fundamental recommendation setting. Instead of isolated modality-specific encoding, OpticalRec renders item textual metadata as visual glyphs, enabling native image-text interaction within the visual encoder - the perceptual encoding level. The resulting representations are further processed by the language decoder - the semantic encoding level, allowing OpticalRec to exploit the dual-attention mechanism of modern vision-language models that previous encoding methods omitted. OpticalRec's efficacy is theoretically supported by mutual information analysis and empirically demonstrated through superior performance across strong baselines and benchmarks. As a plug-and-play module, OpticalRec (1) introduces minimal cost, (2) is robust against rendered text font, color and layout, etc., and (3) integrates seamlessly into existing multimodal collaborative filtering models.