🤖 AI Summary
This work identifies three fundamental clustering failure modes—distribution collapse, similarity distortion, and structural bias—arising from mismatched sliding-window sizes relative to time-series length during preprocessing. Leveraging integrated computational experiments and probabilistic geometric analysis, we establish the first theoretical model of windowed time-series embeddings, rigorously characterizing the existence and precise boundary conditions under which these failures occur. We propose the first systematic explanatory framework that elucidates the structural distortion mechanisms induced by windowing, thereby filling a critical gap in the robustness theory of time-series clustering preprocessing. Empirical validation across synthetic and real-world datasets confirms the reproducibility of all three failure modes and demonstrates their severe, catastrophic impact on clustering performance—often causing abrupt, order-of-magnitude degradation. Our findings provide foundational theoretical insights and practical warnings for robust time-series representation learning.
📝 Abstract
Things can go spectacularly wrong when clustering timeseries data that has been preprocessed with a sliding window. We highlight three surprising failures that emerge depending on how the window size compares with the timeseries length. In addition to computational examples, we present theoretical explanations for each of these failure modes.