🤖 AI Summary
This study addresses whether temporal link prediction necessarily depends on node representation learning. We propose an extremely minimalist, embedding-free baseline that introduces statistical language modeling to this task. Methodologically, the approach leverages transition and co-occurrence counts combined with Kneser-Ney smoothing, performing predictions through log-linear rule composition. The entire model comprises merely nine to thirteen parameters. Experimental results demonstrate that this method achieves the highest mean reciprocal rank on seven out of sixteen datasets, comprehensively outperforming existing baselines such as EdgeBank. This research confirms that simple counting mechanisms can effectively substitute for complex neural memory modules, thereby establishing a concise and efficient new benchmark for temporal link prediction.
📝 Abstract
Many temporal link predictors summarize past interactions through learned node representations. We examine whether simple counts of recurring interaction patterns can provide competitive predictions without learning these representations. We propose a temporal link predictor based on statistical language modelling. It pools transition and co-occurrence counts across sources to predict links that a source has never formed. We smooth sparse estimates using destination frequencies or Kneser-Ney continuation counts. A shared log-linear rule combines these estimates with popularity, source history, and recency, without node embeddings. In our main evaluation, the model achieves the highest MRR among the compared methods on 7 out of 16 datasets from TGB and TGB-Seq. It also outperforms EdgeBank and Base3 on all 16 datasets and the heuristic family on 14. These gains extend to datasets designed to limit repeated edges. With only 9--13 learned parameters, our model provides a simple and competitive baseline for evaluating future neural temporal link predictors.