sid-based user tokenization

Designs and implements pipelines that convert user behavior traces into sequences of discrete user tokens tied to item or semantic identifiers (sids), including methods to align tokens to shared item sids and to ground tokens in item attributes for downstream modeling. Builds and evaluates these tokenizers for robustness and throughput, engineering compact discrete representations and serving pipelines that scale to high-volume / industrial traffic and runtime constraints.

sid-basedusertokenization

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.19
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Existing Semantic-ID (SID) tokenizers lack a unified diagnostic interface, causing mapping flaws—such as coverage gaps, full-code aliasing, and weak semantic prefixes—to remain undetected until downstream training. This work proposes the first systematic diagnostic framework tailored for SID mappings, which enables pre-training analysis by defining an adapter contract that integrates item mappings, metadata, and generation trajectories. The approach decouples addressability from behavioral semantic prefix evaluation and introduces mapping-level probes—including utilization rate, aliasing rate, neighborhood alignment, popularity distribution, and structural cost—as well as dynamic trajectory hooks. Experiments reveal that GRID-style mappings exhibit an aliasing rate as high as 0.977, whereas ReSID and GAOQ show no aliasing; deterministic category prefixes achieve the strongest co-occurrence alignment (0.447), confirming prefix alignment as a viable signal for candidate exposure.

code aliasinggenerative recommendationmapping inspection

This work addresses the semantic degradation inherent in Semantic IDs (SIDs) within generative recommendation systems, where the SID construction process discards fine-grained semantic structures despite preserving coarse-grained item organization and generation constraints. Through systematic analysis of SIDs across encoding, construction, and autoregressive generation stages, the study reveals that while global semantics are retained, local semantic fidelity is compromised, adversely affecting recommendation performance. To mitigate this issue without introducing additional parameters or requiring model retraining, the authors propose Item-Supported Decoding (ISD)—a lightweight inference-time strategy that leverages user-specific ranking to reinforce SID prefix retention and dynamically reorder generated candidates, thereby alleviating the loss of target items during filtering. Evaluated across three Amazon datasets and eight SID configurations, ISD consistently outperforms baselines, achieving up to a 31.2% relative improvement in NDCG@10.

autoregressive generationgenerative recommendationitem representation

This work addresses the limitations of conventional recommendation systems, where fixed-dimensional dense user embeddings exhibit constrained representational capacity, and existing large language model (LLM)-based textual user tokens struggle to align with item attributes and model deep behavioral sequences. To overcome these challenges, the authors propose TokenMinds, a novel system that introduces semantic IDs (SIDs)—discrete representations—into user modeling for the first time. TokenMinds employs a pretrained LLM-based encoder-decoder architecture to jointly generate interpretable SID user tokens and dense embeddings, effectively balancing semantic expressiveness with compatibility for downstream recommendation tasks while enabling unified cross-scenario modeling. The system has been fully deployed across multiple YouTube production environments, demonstrating significant improvements in ranking performance and validating the effectiveness and complementary value of SID tokens in billion-scale industrial recommendation systems.

dense embeddingsdiscrete representationsrecommender systems

Existing generative recommender systems suffer from a disconnect between semantic ID (SID) construction and personalized ranking objectives, which limits retrieval performance. This work proposes DIG, a novel framework that unifies ranking and retrieval through the lens of tokenization for the first time: it embeds a tokenizer within a discriminative ranking model and trains the entire system end-to-end, leveraging user-item cross features to guide codebook boundary optimization. Additionally, a user-to-token (u2t) distillation module is introduced to enable efficient inference. By design, the ranking model inherently acquires retrieval capabilities, leading to significant improvements across ranking, retrieval, and joint tasks on three public benchmarks and two industrial datasets.

discriminative rankinggenerative retrievalpersonalization

Latest Papers

What's happening recently
View more

This work addresses the underrepresentation and path misalignment of cold-start items in generative recommendation systems caused by static semantic ID assignment. To tackle this, the authors propose a three-stage dynamic optimization framework that identifies static ID allocation as a key performance bottleneck. By decoupling tokenization from the generation objective, the framework introduces intent-aware tokenization and counterfactual contrastive learning to construct a behavior-aligned pool of semantic ID candidates. It further incorporates frozen-backbone evaluation without retraining and dynamic weighted beam search to maintain multiple hypotheses and progressively refine semantic IDs. Evaluated on three Amazon benchmarks, the method significantly outperforms existing generative and sequential recommendation approaches, achieving notable improvements on cold-start metrics.

cold-startgenerative recommendationitem retrieval

This work addresses a critical limitation in existing generative recommender systems, which treat semantic IDs as isolated symbols and thereby fail to capture relationships between semantically similar items that share no overlapping IDs. To overcome this, we propose TopoGR, a novel framework that introduces factorizable binary semantic IDs into generative recommendation for the first time, explicitly modeling the topology of Hamming space. TopoGR preserves semantic proximity during both training and inference through topology-aware input encoding, Hamming distance-based metrics, soft-target supervision, and consistency-aware reranking. Extensive experiments on four benchmark datasets demonstrate that TopoGR significantly outperforms state-of-the-art methods, validating the effectiveness of leveraging Hamming-space topology for modeling item relevance beyond exact ID matching.

Generative RecommendationItem RelatednessLatent Structure

This work addresses the challenge of efficiently integrating traditional heterogeneous signals into large-scale Transformer-based recommendation models, which often suffer from excessively long prompts, high memory consumption, and substantial computational overhead. To overcome these limitations, the authors propose Token Factory, a novel framework that introduces a “soft token” mechanism. This approach compresses multi-source heterogeneous features and encodes them into compact, Transformer-compatible representations, enabling direct input to large models without causing prompt length explosion. Evaluated in industrial-scale recommendation scenarios, the method significantly reduces both memory footprint and computational cost while consistently improving recommendation performance.

computational overheadefficient integrationheterogeneous signals

Hot Scholars

MT

Maurizio Tesconi

Head of Cyber Intelligence Lab - IIT - CNR
Social Media AnalysisCyber IntelligenceBig DataText Mining
WQ

Walter Quattrociocchi

Full Professor @Sapienza University of Rome
Data ScienceNetwork ScienceComplex SystemsCollective Dynamics
MC

Matteo Cinelli

Assistant Professor @Sapienza University of Rome
Data ScienceNetwork ScienceSocial MediaComputational Social Science