🤖 AI Summary
This study addresses the interpretability challenge in time-series classification, where discriminative information is often obscured within time-frequency representations. To this end, we propose the XACT framework, which transcends the limitations of fixed transform domains by learning sparse attribution masks through arbitrary invertible time-frequency transforms, such as the short-time Fourier transform and continuous or discrete wavelet transforms. Furthermore, the virtual detection layer for Layer-wise Relevance Propagation (LRP) is extended to wavelet transforms to enable multi-representation explanations. Experiments on both synthetic and real-world datasets demonstrate that the proposed method generates precise and sparse explanations while significantly reducing spurious feature highlighting. Consequently, XACT effectively enhances the transparency and trustworthiness of model decisions, offering a robust approach to explaining time-series classifiers operating in diverse time-frequency domains.
📝 Abstract
Time-series explainability remains challenging because discriminative information is often encoded in latent frequency or time-frequency features rather than in the raw signal itself. Existing attribution methods typically operate either in the time domain or in a fixed transform domain, limiting their ability to capture salient information across different representations. We propose XACT, a general framework that learns sparse attribution masks over coefficients from arbitrary invertible time-frequency transforms. We evaluate the framework on the STFT, the continuous wavelet transform, and the discrete wavelet transform. In addition, we extend the virtual inspection layer approach from the STFT to both wavelet transforms, enabling LRP to generate explanations in these representations. On a synthetic dataset, XACT produces precise explanations and is less prone to highlighting spurious features than the tested baselines. Across two real-world datasets, XACT produces sparse and structured explanations, although no method performs best across all quantitative evaluation criteria. These results demonstrate that learning explanations directly in time-frequency representations offers a flexible approach to interpreting deep-learning models for time series data.