🤖 AI Summary
This study addresses the challenge of aligning video generation models—whose capabilities and costs vary significantly—with personalized user demands. To this end, we propose a cost-aware personalized video generation router that jointly models textual prompts, subjective user preferences, and budget constraints. By leveraging multi-model comparisons and human preference annotations, the routing algorithm achieves an optimal trade-off between generation quality and economic cost. Furthermore, we release a high-quality dataset annotated with user preferences. Experimental results demonstrate that the proposed router is competitive in preference alignment tasks while substantially reducing average generation costs, exhibiting particularly pronounced advantages under strict budget constraints.
📝 Abstract
Rapid progress in video generation has led to a plethora of models that differ substantially in capability and generation cost. This raises a natural question: can each request be efficiently routed to an appropriate model? We find that even when the consensus of the other annotators is used as an oracle, it agrees with each annotator's own favorite only 34-55% of the time. Motivated by this observation, we introduce TasteRoute, a personalized video-generation router that selects a generator jointly based on the input request, user preferences, and available generation budget. Across text-to-video and image-to-video settings, TasteRoute is competitive with strong simple baselines on preference routing while reducing average generation cost. The cost saving increases under higher budget caps. Finally, we release TasteRoute-3k, a human-annotated dataset containing multi-model video comparisons, quality judgments, preference rankings, and user-profile signals to facilitate future research on personalized and cost-aware video routing.