🤖 AI Summary
This study addresses the limited interpretability of existing deep clustering methods, whose latent clusters are difficult to associate with observable waveform characteristics. To this end, we propose WAVE, a framework that pioneers joint representation learning by treating time-series signals and their deterministically rendered waveform images as complementary views. By integrating contrastive learning with multimodal fusion strategies, WAVE achieves deep alignment between fine-grained temporal variations and holistic visual patterns, yielding highly discriminative embeddings that endow clustering results with intuitive, waveform-level traceability. Extensive experiments on ten public datasets demonstrate that WAVE attains state-of-the-art macro-averaged clustering performance and the best average ranking. Furthermore, case studies validate that the discovered cluster structures can be directly inspected through raw waveforms, thereby bridging the gap between latent representations and physically interpretable signal features.
📝 Abstract
Multivariate Time Series (MTS) clustering is an important tool in temporal data mining, aiming to discover latent group structures from complex observations without supervision. Although existing deep clustering methods can learn discriminative temporal representations, the resulting latent clusters are often difficult to relate back to waveform characteristics that practitioners can directly inspect and compare, limiting their ability to assess whether the discovered patterns reflect meaningful temporal behaviors. This paper, therefore, proposes WAVE (Waveform Aligned Visual-temporal Embedding), which treats time series and their deterministically rendered waveform plots as complementary views of the same observations. To produce discriminative representations whose cluster structures can be traced to observable waveform characteristics, WAVE aligns and integrates fine-grained temporal variations with holistic visual patterns, while associating each discovered cluster with its centroid-nearest authentic sample. Accordingly, interpretability in this work specifically refers to waveform-level traceability rather than a general explanation of model decisions. Extensive evaluations across 10 real-world public datasets show that WAVE achieves the highest macro-averaged clustering performance and the best average rank among the compared methods, while qualitative case studies illustrate how the discovered clusters can be inspected through authentic waveform records. The source code is available at https://github.com/Zheng-Zhu1/WAVE.