Score
Designs, implements, and evaluates recommender systems that use generative models to produce candidate items, personalized ranked lists, or natural‑language suggestions; this includes specifying model architectures (e.g., autoregressive, seq2seq, diffusion), training objectives, decoding strategies, and integration with retrieval and reranking pipelines. It also covers analyzing model behavior and outputs for relevance, diversity, calibration, controllability, and biases, and engineering deployment components such as latency- and safety-aware generation.
Traditional discriminative recommender systems suffer from limitations in semantic understanding, multi-step reasoning, and interactive capability. Method: This work reframes recommendation as a generative paradigm, proposing a three-tier technical framework—data augmentation, model alignment, and task execution—that integrates large language models (LLMs) and diffusion models, augmented with knowledge injection, agent simulation, heterogeneous signal unification, and prompt optimization. Contributions: (1) We introduce the first three-dimensional analytical framework that systematically characterizes five key advantages of generative recommendation: world-knowledge integration, multi-step reasoning, creative content generation, cross-domain generalization, and natural interaction. (2) We comprehensively survey data construction paradigms, model taxonomies, and task formulations for generative recommendation. (3) We delineate the technical roadmap and future directions toward intelligent recommendation assistants endowed with understanding, reasoning, and generative capabilities.
This paper addresses the inadequacy of traditional accuracy metrics in evaluating generative recommender systems (Gen-RecSys). We propose the first multidimensional evaluation framework targeting factual consistency, safety, fairness, and user-intent alignment. Methodologically, we systematically categorize Gen-RecSys risks into two classes: *hallucinatory generation* and *bias/privacy leakage*. Our framework adopts a scenario-driven, multi-metric evaluation paradigm, integrating prompt engineering, fact-checking models, bias auditing tools, dialogue safety protocols, and interpretability analysis. We further release the first open-source prototype benchmark for Gen-RecSys evaluation. Experimental results demonstrate that our framework effectively detects item hallucinations, quantifies recommendation bias, and verifies policy compliance—thereby significantly enhancing evaluation comprehensiveness, trustworthiness, and deployment accountability.
Generative recommender systems suffer from semantic misalignment between language models (LMs) and collaborative filtering (CF) spaces, hindering effective feature alignment. To address this, we propose DMRec—a model-agnostic framework that bridges user interactions and LM outputs via a probabilistic meta-network, and introduces a novel three-stage cross-space distribution alignment mechanism optimized using KL and Jensen–Shannon divergences. This is the first approach to achieve joint distribution matching between CF and language spaces. DMRec supports both LM adaptation and generative modeling of implicit feedback, ensuring plug-and-play compatibility and semantic equivalence across representations. Evaluated on three public benchmarks, DMRec consistently enhances the performance of three distinct generative recommendation models, outperforming state-of-the-art LM-augmented methods. Our results empirically validate that explicit distribution alignment significantly improves generative recommendation—demonstrating both effectiveness and broad applicability.
To address high inference latency and weak sequence modeling in multi-stage recommender systems’ re-ranking tasks under real-time industrial settings, this paper proposes NAR4Rec, a non-autoregressive generative re-ranking model. It is the first to introduce the non-autoregressive paradigm into recommendation re-ranking. We design a sequence-level unlikelihood training objective to suppress invalid sequences, incorporate matching-aware auxiliary modeling to enhance intra-list item correlation learning, and propose a contrastive decoding mechanism to improve feasibility discrimination of candidate sequences. The model supports end-to-end joint training and efficient inference. Offline experiments demonstrate significant improvements over state-of-the-art methods. Online A/B tests show statistically significant gains in click-through rate (CTR) and user session duration. NAR4Rec has been fully deployed in the Kuaishou mobile application, serving over 300 million daily active users.
This paper investigates how foundation models (e.g., GPT, LLaMA, CLIP) fundamentally reshape recommendation systems, focusing on three directions: feature enhancement, generative recommendation, and agent-based interaction. It identifies core challenges in deep integration—including representation alignment, generation controllability, and interaction rationality—and proposes FM4RecSys, the first three-dimensional paradigm framework for foundation model–enabled recommender systems. The framework critically compares performance trade-offs and applicability boundaries across the three pathways. By unifying multimodal representation learning, generative modeling, and agent coordination mechanisms, the work synthesizes state-of-the-art advances and distills key open problems: cross-modal semantic gaps, inference efficiency bottlenecks, and the lack of rigorous evaluation protocols. Finally, it delivers a practical technology roadmap for next-generation, foundation model–driven recommendation systems, offering both theoretical foundations and actionable implementation guidance.
Existing research on generative AI–driven multi-objective recommendation lacks a systematic survey, with gaps in theoretical foundations, evaluation protocols, and technical methodologies. Method: We introduce the first taxonomy for generative AI–enhanced multi-objective recommendation, integrating large language models, diffusion models, prompt engineering, multi-task learning, and causal inference. We propose a unified evaluation framework incorporating 12+ benchmark datasets and 20+ metrics, and conduct a comprehensive analysis of over 100 state-of-the-art studies. Contribution/Results: We distill synergies and trade-offs among five core objectives—fairness, explainability, diversity, privacy, and sustainability—and identify five persistent challenges. Finally, we outline actionable future research directions. This work fills a critical survey gap at the intersection of generative AI and multi-objective recommendation, providing both theoretical grounding and practical guidance for the community.
This work addresses the inefficiencies arising from the tight coupling between feature engineering and model architecture, which hinders rapid iteration, complicates deployment, and impedes reusability—particularly in low-latency online serving scenarios. To resolve this, we propose the Prompt Generation (PG) framework, which decouples feature processing logic from the model through two declarative JSON configuration files, thereby unifying offline training and online inference pipelines. PG introduces a novel configuration-driven mechanism for high-order tokenization and feature assembly, leveraging four feature categories, three composable processing components, and built-in sequence compression to establish a standardized, general-purpose inference pipeline. Deployed in Taobao Search, PG has delivered statistically significant gains of +0.47% in transaction count and +0.51% in GMV, and has been adopted by multiple search and recommendation teams as the standard iterative paradigm for generative retrieval.
This study investigates whether generative recommender systems exacerbate filter bubbles, particularly when incorporating semantic ID (SID) sequences compared to traditional approaches. To this end, the authors propose RecLoop, a large language model–based closed-loop simulation framework that models dynamic user–system interactions through iterative feedback loops, and introduce multidimensional evaluation metrics, including “code-space structural filter bubbles.” The findings reveal that generative recommenders exhibit weaker filter bubble effects at the exposure level than conventional sequential models, yet still show concentration tendencies in SID space. Moreover, collaborative-signal tokenization leads to stronger filter bubbles than semantic tokenization, while increasing model scale helps preserve diversity by retaining niche content.
This work addresses the data sparsity challenge in generative recommender systems under user and item cold-start scenarios by establishing the first standardized evaluation protocol for cold-start recommendation. The authors systematically reproduce and analyze state-of-the-art generative recommendation approaches based on pretrained language models (PLMs), carefully controlling key variables such as model scale, identifier design, and training strategies. Through comprehensive ablation and comparative experiments, they reveal that the performance of existing methods is significantly constrained by a confluence of confounding design choices rather than inherent architectural limitations. By providing a rigorous, reproducible benchmark, this study enhances the reliability of empirical conclusions in cold-start recommendation research and offers clear directions for future improvements.
This work addresses a fundamental limitation in autoregressive generative recommender models, where the tree-based decoding structure of semantic IDs induces coupling among item probabilities, hindering the model’s ability to capture fine-grained user preferences between neighboring items. The study is the first to formally reveal how this hierarchical decoding architecture inherently constrains model expressiveness. To overcome this issue, the authors propose Latte, a novel approach that injects learnable latent tokens before each semantic ID, effectively decomposing the single decoding tree into multiple conditional subtrees and thereby decoupling item generation probabilities. Extensive experiments demonstrate that Latte consistently improves recommendation accuracy, yielding an average gain of 3.45% in NDCG@10 across benchmark datasets.