Recommendation World Models for Future-State Control

πŸ“… 2026-09-24
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the challenge in sequential recommendation where dynamically presented content influences user states and future feedback, complicating long-term decision optimization. To this end, this work proposes the UA-TWM framework, which achieves objective-aware decision-making centered around a pre-trained ranker. By constructing neighboring actions and evaluating their local consequences, the method selects optimal alternatives under utility constraints to regulate future states. Furthermore, it introduces an innovative utility anchoring mechanism that integrates log replay, closed-loop state-action prediction, and calibrated failure risk estimation to ensure decision reliability. Extensive experiments on multiple datasets demonstrate that UA-TWM significantly improves Recall@20 and NDCG@20, effectively enhancing the system’s capacity for precise alignment with future states.
πŸ“ Abstract
Sequential recommendation optimizes which items to rank, while each displayed slate also shapes subsequent feedback and user state. We study how a trained ranker can support decisions about these future consequences. We introduce UA-TWM, a utility-anchored world-model interface that constructs nearby slate actions, estimates their target-relevant consequences, and selects an alternative subject to utility constraints. The reference slate serves as a fallback when no alternative qualifies. A logged-replay instantiation combines utility and target-gain estimates with calibrated failure-risk prediction; a closed-loop instantiation uses one-step state-action prediction and updates its decisions after observed feedback. We evaluate transfer across twelve sequential backbones on MovieLens-25M and KuaiRand-Pure, and repeated target-directed interaction in KuaiSim. Attaching the interface improves Recall@20, NDCG@20, and future-state alignment for every matched logged backbone. Selection ablations reveal the utility and risk costs of aggressive target pursuit, while closed-loop diagnostics isolate the contribution of action-conditioned prediction. Local consequence modeling thus enables target-aware selection around a trained sequential ranker.
Problem

Research questions and friction points this paper is trying to address.

Sequential Recommendation
World Models
Future-State Control
Slate Optimization
Innovation

Methods, ideas, or system contributions that make the work stand out.

World Models
Sequential Recommendation
Future-State Control
Utility Constraints
Closed-loop Prediction