Score
Designs and evaluates strategies for segmenting temporal data into analysis windows by choosing window length, overlap (overlapping vs non‑overlapping), and sampling density. Builds and analyzes windowing schemes to capture relevant temporal dynamics while minimizing redundancy and preserving generalization and statistical power.
This work identifies three fundamental clustering failure modes—distribution collapse, similarity distortion, and structural bias—arising from mismatched sliding-window sizes relative to time-series length during preprocessing. Leveraging integrated computational experiments and probabilistic geometric analysis, we establish the first theoretical model of windowed time-series embeddings, rigorously characterizing the existence and precise boundary conditions under which these failures occur. We propose the first systematic explanatory framework that elucidates the structural distortion mechanisms induced by windowing, thereby filling a critical gap in the robustness theory of time-series clustering preprocessing. Empirical validation across synthetic and real-world datasets confirms the reproducibility of all three failure modes and demonstrates their severe, catastrophic impact on clustering performance—often causing abrupt, order-of-magnitude degradation. Our findings provide foundational theoretical insights and practical warnings for robust time-series representation learning.
In streaming process mining, fixed-size sliding windows struggle to accommodate dynamic process evolution and concept drift, leading to model bias. To address this, we propose a dynamic window optimization method grounded in species estimation theory—introducing, for the first time, sample representativeness quantification into streaming process mining. Our approach establishes a real-time representativeness assessment model under sliding windows and adaptively adjusts window size to balance timeliness and statistical sufficiency. It requires no prior knowledge and enables online detection and response to concept drift. Experiments on multiple real-world event streams demonstrate that our method significantly improves process model accuracy (average +12.7% F1-score) and robustness to concept drift (38.5% reduction in false positive rate) compared to static-window baselines. This work establishes a novel paradigm for real-time, adaptive process analysis.
This paper addresses the challenging problem of clustering multivariate, irregularly sampled time series with missing observations—common in healthcare and psychology. We propose the first mixture-clustering framework based on linear Gaussian state-space models (LGSSMs). Our method explicitly embeds parameterized LGSSMs into a mixture model, derives a closed-form log-likelihood function, and designs an expectation-maximization (EM) algorithm tailored to irregular sampling, thereby ensuring both theoretical interpretability and empirical robustness. The approach naturally accommodates varying sampling rates and models missing data via the underlying probabilistic structure. Extensive evaluation on synthetic data and real-world applications—including wearable sensing, ecological momentary assessment, and epidemiological cohort studies—demonstrates significant improvements over existing time-series clustering methods. The implementation is publicly available.
This study addresses the limitations of conventional temporal segmentation in capturing the dynamic evolution of learning behaviors and identifying intervention windows by proposing a Topological Data Analysis (TDA)-based approach to learning cycle segmentation. Specifically, the method leverages zigzag persistent homology to extract change points in connected components (β0) from transition networks, enabling data-driven segmentation. Its core contribution is reconceptualized as detecting structural behavioral shifts to pinpoint timely interventions, rather than merely enhancing predictive accuracy. Experiments conducted across 22 courses at an open university demonstrate that the proposed approach significantly outperforms traditional time-based baselines. It captures more pronounced inter-period variations and effectively reveals the association between learner behavioral collapse and academic outcomes.
In model-based clustering, incomplete prior knowledge and difficulty in specifying appropriate uncertainty over the entire partition hinder robust inference. Method: We propose a locally weighted probabilistic modeling framework that allocates prior uncertainty at the individual observation level—departing from conventional global penalty schemes—and integrates spatiotemporal dependence via a spatiotemporal Gaussian process, coupled with a Bayesian nonparametric random partition model and MCMC inference. Contribution/Results: Our approach enables fine-grained, subset-specific uncertainty encoding, substantially enhancing flexibility and robustness in incorporating expert knowledge. Evaluated on PM₁₀ spatiotemporal data and synthetic experiments, it achieves 12–19% higher clustering accuracy than baseline methods and demonstrates superior robustness to misspecified priors.
This study addresses the representation, visualization, and quantification of regular structures in rhythmic data by proposing a “rhythmic motif” framework that models rhythms as fixed-length sequences of inter-onset intervals, decomposed into duration and pattern (the ratios between successive intervals). Building on this decomposition, the authors introduce pattern–duration plots and cluster transition networks to uncover rhythmic regularities. They reformulate the normalized Pairwise Variability Index (nPVI) as the mean deviation from isochrony and propose a more general measure of anisochrony alongside a novel concept termed “quantization.” This approach unifies existing visualization techniques and effectively reveals small-integer-ratio rhythmic structures in both synthetic and real-world datasets, establishing a more flexible and theoretically grounded framework for rhythm analysis.
This study addresses the lack of design guidelines for highlighting current values in small multiples time-series visualizations. Grounded in visual channel theory and iterative design, this work constructs a design space that integrates current and historical data, validating its effectiveness through an online empirical experiment. The findings demonstrate that integrated designs outperform separated layouts, while establishing best practices for size and color encoding in threshold-based tasks. Specifically, the integrated approach improves response times by 28% for non-threshold tasks, and color encoding significantly enhances threshold detection performance. These results provide a theoretical foundation and practical design specifications for facilitating rapid assessment in real-time monitoring scenarios.
该研究针对选区重划中的高维组合问题,通过构建和诊断具有适当覆盖和重叠的候选层来实现分层抽样,使用聚类方法形成计划层面的'单词'。
This study systematically investigates how uncertainty estimation in segmentation tasks is influenced by multiple factors, including dataset difficulty, model architecture, and temporal information. We introduce the first multivariate evaluation framework tailored for semantic and panoptic segmentation, comprehensively analyzing the impact of different backbone networks, datasets, and downstream task configurations on the quality of uncertainty estimates derived from deterministic models, ensemble methods, and temporal ensembles. Our experiments reveal that uncertainty calibration is notably more challenging in panoptic segmentation; temporal information improves calibration only under specific conditions; while sample diversity offers limited gains over simple baselines, ensemble methods—when properly deployed—consistently outperform deterministic approaches. This work establishes a systematic benchmark and provides practical guidance for uncertainty modeling in segmentation tasks.
TSExplorer通过多角度2D可视化方法,解决了时间序列数据的交互式标注与探索问题。