Unpaired Modality-Agnostic Generative Recommendation

📅 2026-08-03
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limitations of existing multimodal generative recommendation methods, which rely on paired data and are constrained by modality-specific semantic ID spaces, thereby struggling to effectively leverage unpaired unimodal observations. To overcome this, the authors propose a unified semantic ID space framework that enables modality-agnostic generative recommendation through a shared Transformer encoder, a residual vector quantization codebook, and lightweight modality-specific projection heads. This design seamlessly integrates paired, image-only, and text-only data without requiring feature imputation or multiple codebooks. Extensive experiments on three benchmark datasets demonstrate that the proposed method significantly outperforms current state-of-the-art models under both complete and missing modality settings, confirming its strong cross-modal generalization capability and practical utility.
📝 Abstract
Generative Recommendation (GR) formulates recommendation as autoregressive generation over discrete semantic identifiers (IDs). Although recent multimodal GR methods improve semantic ID construction with visual and textual information, they typically require item-level paired observations, restricting tokenization to the intersection of modality availability. Moreover, incorporating unpaired observations is nontrivial because small representation shifts may cross quantization boundaries and produce incompatible identifier sequences. To address this challenge, we propose \textbf{Unpair}ed Modality-Agnostic \textbf{G}enerative \textbf{R}ecommendation (UnpairGR), which learns a unified semantic-ID space from paired, image-only, and text-only observations. UnpairGR confines modality-specific processing to lightweight input projections while sharing the subsequent Transformer and residual codebooks across all observation conditions. Paired observations establish a reliability-guided cross-modal consensus, whereas unimodal observations directly refine the same representations and codes. The learned tokenizer is then fixed to provide stationary targets for a single autoregressive recommender, without feature imputation, modality-specific codebooks, or fallback mappings. Extensive experiments on three benchmark datasets demonstrate that UnpairGR consistently improves recommendation performance under both fully observed and incomplete-observation settings.
Problem

Research questions and friction points this paper is trying to address.

Generative Recommendation
Unpaired Multimodal Data
Semantic Tokenization
Modality-Agnostic Learning
Quantization Boundaries
Innovation

Methods, ideas, or system contributions that make the work stand out.

Generative Recommendation
Modality-Agnostic
Unpaired Multimodal Learning
Semantic Tokenization
Unified Representation
🔎 Similar Papers
No similar papers found.
W
Weihao Shen
Institute of Artificial Intelligence, Beihang University, Beijing, China
W
Wei Chen
Institute of Artificial Intelligence, Beihang University, Beijing, China
F
Fuwei Zhang
Institute of Artificial Intelligence, Beihang University, Beijing, China
Meng Yuan
Meng Yuan
Marie Skłodowska-Curie Fellow, Chalmers University of Technology
MechatronicsEnergy systemModel predictive controlRobotics
Y
Yuqin Lan
Institute of Artificial Intelligence, Beihang University, Beijing, China
G
Guojun Liu
Meituan, Beijing, China
Q
Qingsong Hua
Meituan, Beijing, China
W
Wei Lin
Meituan, Beijing, China
F
Fuzhen Zhuang
Institute of Artificial Intelligence, Beihang University, Beijing, China