A Systematic Benchmark of Explainable Methods for Temporal Attribution in Sequential Recommendation Systems

📅 2026-09-22
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
本文通过引入双模型掩码度量方法,系统地评估了十种可解释性方法在序列推荐系统中的归因准确性,解决了深度学习模型决策难以解释的问题。
📝 Abstract
Sequential RecSys are central to modern personalization, exploiting user's historical interaction sequences to drive next-step decisions. Deep learning models, particularly CNN and Transformer-based architectures, have proven highly effective at capturing temporal dependencies in these histories. For transparency and trust, understanding which past interactions drive a given recommendation is increasingly important --- both for developers auditing model behavior and for users seeking a rationale. However, the non-linearities that give these models their predictive power also render them black boxes, making it difficult to attribute decisions to specific interactions. While gradient-based, perturbation-based, and attention-based explainability methods exist, a systematic benchmark of their faithfulness for sequential recommendation is missing. We address this gap by introducing a dual-model masking metric in which one model supplies per-timestep attribution scores and a separately trained, masking-robust probe measures the resulting change in predicted probability. Using this metric, we benchmark ten XAI methods across CNN, Transformer, SASRec, and BERT4Rec backbones on KuaiRand and MovieLens, complemented by analyses of temporal attribution patterns, item popularity confounding, and robustness to input corruption. Our key findings are: (1) gradient-based methods, particularly GradientSHAP and Integrated Gradients, yield the most faithful and robust attributions; (2) raw attention weights are unreliable, but gradient-weighted attention restores faithfulness on shorter sequences, with degradation on longer horizons as softmax attention probabilities converge toward uniform importance scores, diminishing the method's ability to identify informative interactions; and (3) temporal attribution patterns in faithful methods reflect genuine task structure rather than recency or popularity bias.
Problem

Research questions and friction points this paper is trying to address.

Sequential Recommendation
Explainability
Temporal Attribution
Deep Learning Models
Transparency
Innovation

Methods, ideas, or system contributions that make the work stand out.

Dual-model masking metric
Gradient-based methods
Temporal attribution
Sequential recommendation systems
XAI
Akash Pandey
Akash Pandey
Northwestern University
ML for proteinExplainable AIDeep-LearningFinite Element Analysis
K
Kanisha Shah
Capital One, AI Foundations
A
Addrish Roy
Capital One, AI Foundations
D
Dwipam Katariya
Capital One, AI Foundations
H
Hongyangyang Shi
Capital One, AI Foundations
A
Amanda Ding
Capital One, AI Foundations
K
Kalanand Mishra
Capital One, AI Foundations
Pranab Mohanty
Pranab Mohanty
Capital One
Generative AILLMSafe AIRecommender SystemDeep Learning