🤖 AI Summary
This paper addresses the clustering of functional data residing in infinite-dimensional Hilbert spaces—a long-standing challenge in functional data analysis. We introduce the first mean shift algorithm adapted to the functional domain, accompanied by a rigorous theoretical framework establishing its convergence and stability. To overcome computational bottlenecks inherent in large-scale functional datasets, we propose a scalable stochastic variant based on random partitioning and block-wise iteration, which retains strong theoretical guarantees while achieving substantial efficiency gains. The method integrates functional data analysis, kernel density estimation, and randomized subsampling techniques. Empirical evaluation on Argo ocean profile data demonstrates both effectiveness and scalability: the stochastic version achieves high-fidelity approximation of the original algorithm while drastically reducing computational cost, thereby significantly enhancing the practicality and efficiency of clustering massive functional datasets.
📝 Abstract
This paper extends the mean shift algorithm from vector-valued data to functional data, enabling effective clustering in infinite-dimensional settings. To address the computational challenges posed by large-scale datasets, we introduce a fast stochastic variant that significantly reduces computational complexity. We provide a rigorous analysis of convergence and stability for the full functional mean shift procedure, establishing theoretical guarantees for its behavior. For the stochastic variant, although a full convergence theory remains open, we offer partial justification for its use by showing that it approximates the full algorithm well when the subset size is large. The proposed method is further validated through a real-data application to Argo oceanographic profiles. Our key contributions include: (1) a novel extension of mean shift to functional data; (2) convergence and stability analysis of the full functional mean shift algorithm in Hilbert space; (3) a scalable stochastic variant based on random partitioning, with partial theoretical justification; and (4) a real-data application demonstrating the method's scalability and practical usefulness.