🤖 AI Summary
This study addresses the high computational overhead, inefficient fixed-rate processing, and noise susceptibility of Transformer-based recommendation models in long-sequence modeling. Inspired by the Byte Latent Transformer, we propose DP-Rec, a dynamic patching architecture that leverages contrastive entropy to detect behavioral boundaries and dynamically compresses item-level sequences into latent patch vectors. This enables a paradigm shift from item-level to patch-level modeling, effectively mitigating computational redundancy and context loss during inference. Experimental results demonstrate that DP-Rec significantly enhances long-sequence scalability under constrained computational budgets, achieving superior efficiency-accuracy trade-offs compared to both uncompressed and fixed-compression baselines.
📝 Abstract
Transformers have redefined sequential recommendation by effectively modeling dynamic user behaviors and long-range dependencies. However, they remain inherently inefficient: standard architectures operate at a fixed rate, allocating comparable computation to every item in a user’s history regardless of its information content. This leads to prohibitive computational overhead on long sequences and increased sensitivity to behavioral noise. To address this, practitioners often resort to lossy sequence compression, staged modeling, or truncation. This limits the model’s ability to leverage the full context of long histories during inference. Inspired by the recent success of Byte Latent Transformer, we propose DP-Rec, a dynamic latent patching architecture for recommendation. DP-Rec shifts from item-level modeling to patch-level modeling by segmenting interaction sequences using contrastive entropy surprise to identify informative behavioral boundaries. A lightweight patch encoder compresses these temporally contextualized segments into a reduced set of dynamic latent behavior vectors, which are then processed by a larger latent transformer and decoded for next-item prediction. Extensive experiments show that, under constrained computational budgets, DP-Rec scales effectively to long sequences and achieves a superior efficiency–accuracy trade-off over both non-compressed and fixed-size compression baselines.