Routing Between Generative and Collaborative User Profiles: A Serving-Time Gate for Controllable Novelty

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the high cost and deployment challenges of large language model (LLM) user profiling, alongside the lack of efficient routing mechanisms between generative and collaborative filtering models in recommender systems. We propose a serving-time-only gating network that dynamically routes users to either collaborative sequential recommendation or LLM-based profiling models based on real-time context, enabling a controllable trade-off between novelty and relevance. Our analysis reveals that performance gains stem from precise selective routing rather than LLM generation alone, supported by an adjustable threshold control strategy. Experiments demonstrate that, under a 5% NDCG degradation constraint, the proposed method accurately routes only 12.5% of users to the LLM while improving Novelty@10 by 6.5%, significantly outperforming heuristic and random baselines.
📝 Abstract
Large language models (LLMs) enable rich semantic user profiles for recommendation, but such profiles are more expensive to generate and are not necessarily desirable to deploy uniformly. We study whether LLM-generated profiles can instead be invoked selectively within a production recommendation pipeline. Using a real-world streaming dataset covering movies, TV shows, and sports content, we train a serving-time routing gate that assigns each user to either a collaborative sequential recommendation model or a recommendation model driven by an LLM-generated profile. The gate uses only serving-time features and learns to identify users for whom profile-based routing can increase Novelty@10 while preserving ranking relevance. A routing threshold controls how aggressively users are sent to the generative model, exposing a tunable novelty--relevance trade-off. At an overall NDCG-loss budget of 5\%, the learned gate increases Novelty@10 by 6.5\% while routing 12.5\% of users, outperforming simple heuristic and random routing policies at comparable relevance cost. These results show that LLM-generated user profiles can serve as a controllable complement to collaborative recommendation, while results with non-generative semantic profiles indicate that the benefit stems from selective routing rather than LLM generation alone.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
User Profiles
Recommendation Systems
Novelty
Routing
Innovation

Methods, ideas, or system contributions that make the work stand out.

Serving-time Routing Gate
LLM-generated User Profiles
Controllable Novelty
Recommendation Systems
Novelty-Relevance Trade-off
🔎 Similar Papers
No similar papers found.