🤖 AI Summary
Addressing the challenges of real-time changepoint detection in semi-structured time-series data (e.g., sensor and video streams)—namely high latency, low accuracy, and difficulty in detecting abrupt disorders—this paper proposes an end-to-end differentiable changepoint detection (CPD) framework. We design a principled, differentiable CPD loss function that jointly optimizes detection latency and false positive rate, integrated with deep representation learning for unified optimization. To support rigorous evaluation, we introduce and publicly release the first video benchmark dataset featuring precise, frame-level disorder annotations. On explosion event detection in videos, our method achieves an F1 score of 0.53, substantially outperforming state-of-the-art baselines (0.31 and 0.35). Furthermore, comprehensive experiments on synthetic sequences and real-world sensor data demonstrate strong generalization and robustness across diverse modalities and noise conditions.
📝 Abstract
For sequential data, a change point is a moment of abrupt regime switch in data streams. Such changes appear in different scenarios, including simpler data from sensors and more challenging video surveillance data. We need to detect disorders as fast as possible. Classic approaches for change point detection (CPD) might underperform for semi-structured sequential data because they cannot process its structure without a proper representation. We propose a principled loss function that balances change detection delay and time to a false alarm. It approximates classic rigorous solutions but is differentiable and allows representation learning for deep models. We consider synthetic sequences, real-world data sensors and videos with change points. We carefully labelled available video data with change point moments and released it for the first time. Experiments suggest that complex data require meaningful representations tailored for the specificity of the CPD task --- and our approach provides them outperforming considered baselines. For example, for explosion detection in video, the F1 score for our method is 0.53 compared to baseline scores of 0.31 and 0.35.