spatio-temporal clustering

Design and implement clustering algorithms that group data across space and time by operating on local 3D windows (e.g., contiguous spatial patches over successive frames or volumetric neighborhoods), restricting centroid search to nearby tokens to exploit spatio-temporal locality. Build low-overhead, per-step clustering updates and parallel-friendly implementations that reduce computation and memory by limiting operations to local neighborhoods for efficient processing of video or other spatio-temporal sequences.

spatio-temporalclustering

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.07
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This paper studies dynamic correlation clustering: given $n$ objects and pairwise similarity/dissimilarity labels, the goal is to partition them into clusters minimizing label disagreements; in the fully dynamic setting, edge labels may flip in real time. We break the long-standing $3$-approximation barrier for this problem by proposing Modified Pivot—a novel algorithm built upon the Pivot framework, incorporating local vertex reassignment and cluster-structure maintenance mechanisms, supported by efficient data structures. It achieves a strict theoretical approximation ratio better than $3$ (e.g., $2.95$). Each update requires only $O(mathrm{polylog},n)$ time, thus simultaneously ensuring high solution quality and low latency. Our approach establishes a breakthrough trade-off between dynamic responsiveness and approximation accuracy.

Breaking 3-approximation barrier in fully dynamic settingDynamic correlation clustering with label flipsMaintaining improved clusters in polylogarithmic update time

On the clustering behavior of sliding windows

Mar 18, 2025
BA
B. Alexeev
🏛️ The Ohio State University | San Jose State University

This work identifies three fundamental clustering failure modes—distribution collapse, similarity distortion, and structural bias—arising from mismatched sliding-window sizes relative to time-series length during preprocessing. Leveraging integrated computational experiments and probabilistic geometric analysis, we establish the first theoretical model of windowed time-series embeddings, rigorously characterizing the existence and precise boundary conditions under which these failures occur. We propose the first systematic explanatory framework that elucidates the structural distortion mechanisms induced by windowing, thereby filling a critical gap in the robustness theory of time-series clustering preprocessing. Empirical validation across synthetic and real-world datasets confirms the reproducibility of all three failure modes and demonstrates their severe, catastrophic impact on clustering performance—often causing abrupt, order-of-magnitude degradation. Our findings provide foundational theoretical insights and practical warnings for robust time-series representation learning.

Clustering failures in sliding window timeseries dataImpact of window size relative to timeseries lengthTheoretical and computational analysis of failure modes

Depth-Based Local Center Clustering: A Framework for Handling Different Clustering Scenarios

May 14, 2025
SW
Siyi Wang
🏛️ McMaster University | University of Manitoba

Existing clustering methods (e.g., K-means, DBSCAN) struggle with multimodal, non-convex, nested, and noisy data, often relying on predefined parameters or global structural assumptions. To address these limitations, we propose Deep-driven Local Centroid Clustering (DLCC). DLCC introduces the novel concept of *local data depth*, overcoming the inadequacy of global depth in characterizing multimodal structures; designs a density-sensitive intra-cluster validity metric to assess internal quality of non-convex clusters; and incorporates subset-based depth ranking with a parameter-free, adaptive mechanism for local centroid identification. Extensive experiments demonstrate that DLCC significantly outperforms state-of-the-art methods on datasets featuring diverse shapes, non-convexity, nesting, and noise. Crucially, DLCC requires no prior specification of the number of clusters, exhibits strong generalizability, and maintains robustness and practical utility across heterogeneous scenarios.

Capturing multimodal data characteristics effectivelyEvaluating performance on non-convex clustersHandling diverse clustering scenarios with limitations

Distributed clustering in partially overlapping feature spaces

Oct 10, 2025
AM
Alessio Maritan
🏛️ University of Padova

This paper addresses distributed clustering under partial feature space overlap, where multiple parties hold private, heterogeneous datasets with partially overlapping features (e.g., cross-institutional healthcare data) and cannot share raw data or full feature sets. To tackle this, we propose two novel federated clustering algorithms: (i) a global centroid aggregation scheme that enables federated updates of shared cluster centers, and (ii) a statistical modeling approach that generates and aggregates synthetic proxy data to align heterogeneous feature spaces. Both methods support participant autonomy in selecting local clustering models and customizing computational overhead. Under mild regularity conditions, the algorithms converge to the centralized optimal solution. Experiments on three public benchmark datasets demonstrate that our methods achieve clustering performance close to the centralized oracle baseline and significantly outperform existing distributed clustering baselines, confirming their practical deployability.

Clustering data distributed across multiple institutions with overlapping featuresDeveloping federated and one-shot algorithms for distributed clustering scenariosSolving feature space heterogeneity when each site has partial feature sets

Fast, Space-Optimal Streaming Algorithms for Clustering and Subspace Embeddings

Apr 22, 2025
VC
Vincent Cohen-Addad
🏛️ Google Research | Texas A&M University | Carnegie Mellon University

This paper investigates asymptotically optimal streaming algorithms for ((k,z))-clustering and subspace embedding. To overcome the fundamental bottlenecks of classical methods—where memory and update time scale with input size (n) and data range (Delta)—we propose the first (n)- and (Delta)-independent streaming framework, integrating core-set compression, hierarchical sampling, random projection, and adaptive reweighting, along with novel (p)-norm sensitivity analysis and dynamic maintenance mechanisms. Our theoretical contributions are: (i) for ((k,z))-clustering, space complexity improves to ( ilde{O}(dk / min{varepsilon^4, varepsilon^{z+2}})) and update time to (d cdot log k cdot mathrm{polylog}(log(nDelta))); (ii) for subspace embedding, we achieve (O(d)) update time and ( ilde{O}(d^2/varepsilon^2)) space—the first streaming algorithm matching the time and space efficiency of offline counterparts.

Achieve space-optimal subspace embeddings in streamingMatch offline algorithm efficiency in streaming modelsOptimize streaming clustering with minimal memory usage

Latest Papers

What's happening recently
View more

This work addresses the node clustering problem in edge-colored hypergraphs, where the goal is to assign colors to nodes so as to maximize agreement with the colors of their incident hyperedges—a problem known to be NP-hard. The paper proposes a novel purely combinatorial approximation algorithm that, for the first time, achieves an approximation factor strictly below 2 without relying on linear programming, thereby breaking through the performance barrier of existing combinatorial approaches. By integrating local search with greedy strategies, the method significantly enhances approximation quality while preserving scalability, offering an efficient and theoretically superior solution to this challenging optimization problem.

approximation algorithmcombinatorial algorithmedge-colored clustering

This work addresses the limitations of traditional clustering algorithms—specifically, the high computational complexity (O(N²)) of density-based methods like DBSCAN and the inability of partitioning approaches such as K-Means to capture nonlinear structures or handle noise effectively. To overcome these challenges, the authors propose K-SCAN, a novel algorithm that synergistically integrates vector quantization with density-based analysis. K-SCAN first employs stochastic mini-batch K-Means to generate weighted micro-clusters and then performs density connectivity analysis on these micro-clusters. This approach achieves linear time complexity while accurately identifying nonlinear manifold structures and exhibiting robustness to noise. Experimental results demonstrate that K-SCAN runs over three times faster than BIRCH on million-scale datasets, attains an Adjusted Rand Index exceeding 0.99, and remains effective even with noise levels as high as 55%.

density-based clusteringnoise robustnessnon-linear clusters

Traditional clustering methods struggle with high-dimensional data where different subsets of features correspond to distinct cluster structures and some features are uninformative. This work proposes a local spectral clustering framework that formulates local clustering as a “clustering of clusterings” problem. It groups features via a label-invariant clustering matrix and constructs feature-specific Gaussian kernel similarity matrices based on a heterogeneous sub-Gaussian mixture model. The approach jointly identifies homogeneous feature groups and their corresponding sample partitions without requiring explicit likelihood evaluation or Bayesian inference. Experiments demonstrate that the method effectively uncovers complex heterogeneous clustering structures in both synthetic and real-world datasets, exhibiting superior performance and practical utility.

clustering-of-clusteringsfeature groupingheterogeneous clustering

This work addresses the challenge of real-time detection of sparse, minute event clusters in event camera data by proposing an asynchronous, event-driven hierarchical agglomerative clustering algorithm. The method leverages a spatiotemporal distance metric between events and triggers clustering immediately upon each event’s arrival, eliminating reliance on frame-based structures or pixel-array dimensions. With linear time complexity O(n), the algorithm achieves both computational efficiency and implementation simplicity, significantly outperforming existing approaches in terms of processing speed and resource consumption. This enables low-latency, high-precision detection of small and sparse event clusters, making it particularly suitable for real-time applications requiring responsiveness and minimal overhead.

asynchronous dataevent cameraevent clustering

Hot Scholars

SF

Shihong Fan

Hyundai America Technical Center Inc.
automotive
DK

Dominik Karbowski

Principal Research Engineer, Argonne National Laboratory
EnergyIntelligent Transportation SystemsControl TheoryOptimization
GF

Guido Fioretti

Università di Bologna, Department of Management Science
Organization ScienceDecision Theory
CL

Chun-Liang Li

Apple / University of Washington
Machine LearningStatistics
JH

Jiaming Han

PhD Student, CUHK MMLab
Computer VisionVision-LanguageVisual Generation