Reinforcement Learning with Conformal Action Sets: An Application to Sequential Recommendation

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of fixed candidate sets in adapting to dynamic session changes within sequential recommendation by proposing RLCP. This method dynamically prunes the action space using critic scores and an online threshold, which is updated via binary feedback to balance diversity and reward. The core innovation lies in introducing a conformal pruning mechanism that establishes deterministic bounds on surrogate missing rates and provides an exact decomposition theory for value loss. Experimental results demonstrate that RLCP achieves superior catalog diversity across multiple benchmarks—ranging from 1.11 to 5.21 times that of the strongest baseline—while maintaining competitive session depth.
📝 Abstract
Sequential recommenders typically use a fixed slate size even though the number of useful alternatives changes within a session. We propose Reinforcement Learning with Calibrated Pruning (RLCP), which adapts the retained action set using critic scores and an online threshold. The threshold is updated from binary feedback indicating whether the set contains an action in a proxy target. We prove a deterministic bound on the observed proxy miss rate along adaptive trajectories. To quantify the effect of pruning on reward, we derive an exact decomposition of value loss into filtering and selection losses. Under explicit proxy and critic approximation conditions, this decomposition yields a finite session reward bound that also accounts for imperfect selection and set truncation, without requiring the learning parameters to converge. Experiments on KuaiRand-Pure and MovieLens 1M compare two RLCP implementations with four RL baselines. In each of the 19 configurations, at least one RLCP variant achieves the highest catalog diversity, reaching $1.11\times$ to $5.21\times$ that of the strongest baseline, with competitive session depth and no larger retained sets.
Problem

Research questions and friction points this paper is trying to address.

Sequential Recommendation
Reinforcement Learning
Adaptive Action Set
Slate Size
Catalog Diversity
Innovation

Methods, ideas, or system contributions that make the work stand out.

Reinforcement Learning
Conformal Action Sets
Sequential Recommendation
Calibrated Pruning
Value Decomposition