arXiv · 2609.39043
Routing Between Generative and Collaborative User Profiles: A Serving-Time Gate for Controllable Novelty
Abstract
Large language models (LLMs) enable rich semantic user profiles for recommendation, but such profiles are more expensive to generate and are not necessarily desirable to deploy uniformly. We study whether LLM-generated profiles can instead be invoked selectively within a production recommendation pipeline. Using a real-world streaming dataset covering movies, TV shows, and sports content, we train a serving-time routing gate that assigns each user to either a collaborative sequential recommendation model or a recommendation model driven by an LLM-generated profile. The gate uses only serving-time features and learns to identify users for whom profile-based routing can increase Novelty@10 while preserving ranking relevance. A routing threshold controls how aggressively users are sent to the generative model, exposing a tunable novelty--relevance trade-off. At an overall NDCG-loss budget of 5\%, the learned gate increases Novelty@10 by 6.5\% while routing 12.5\% of users, outperforming simple heuristic and random routing policies at comparable relevance cost. These results show that LLM-generated user profiles can serve as a controllable complement to collaborative recommendation, while results with non-generative semantic profiles indicate that the benefit stems from selective routing rather than LLM generation alone.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Milad Sabouri, Neeraj Sharma, Sardar Hamidian, Shaghayegh Agah. 2026-09-30. Routing Between Generative and Collaborative User Profiles: A Serving-Time Gate for Controllable Novelty. https://arxiv.org/abs/2609.39043
Cite the original work for its findings. Save a collection to share your selection of sources.