Score
Design and implement clustering algorithms and components that use optimal transport to compute couplings or transport plans between distributions of representations, enabling transport-based matching and alignment of cluster assignments across views or batches. Build objective terms, matching procedures, or regularizers based on transport costs to guide self-supervised cluster formation, enforce consistency, and improve clustering robustness.
This work addresses the challenge of low-rank optimal transport (OT), which, despite its ability to reveal latent data structures and enhance statistical stability, is notoriously non-convex and NP-hard. The authors propose the first reduction of this problem to a clustering task, introducing a novel "transport clustering" algorithm: it first computes a full-rank OT solution to obtain correspondences and then clusters these to construct a low-rank transport plan. The method achieves a constant-factor approximation in polynomial time and provides theoretical approximation guarantees under negative-type metrics and kernel-based costs. Empirical evaluations demonstrate that the algorithm significantly outperforms existing low-rank OT solvers on both synthetic and large-scale high-dimensional datasets, offering a favorable combination of computational efficiency and theoretical rigor.
This work addresses the challenge of robust region alignment in point cloud matching when the data exhibit intrinsic clustering structures. Conventional methods often enforce strict point-to-point correspondences, disregarding the interchangeability of points within clusters and thus failing to preserve regional coherence. To overcome this limitation, the authors propose a clustering-aware matching framework that integrates a quadratic Laplacian regularizer—constructed from a similarity graph—into the optimal transport formulation (termed LapOT), thereby encouraging matchings that respect the underlying cluster structure. Furthermore, they introduce a Refinement via Synchronized Clustering (RSC) mechanism to achieve consistent partitioning across point sets. This approach is the first to incorporate Laplacian regularization into optimal transport for modeling clustering priors, effectively mitigating the drawbacks of independent clustering. Theoretical analysis and experiments demonstrate that the proposed method significantly outperforms existing baselines in preserving cluster integrity and enhancing matching robustness, yielding more consistent and interpretable region alignments.
Conventional dimensionality reduction and clustering methods for high-dimensional data are often decoupled, hindering effective modeling of multi-scale structural patterns. Method: This paper proposes a unified framework based on distributional simplification, which— for the first time—embeds both tasks within the Gromov–Wasserstein (GW) optimal transport geometry. By modeling the intrinsic metric structure via GW projection, the framework jointly learns low-dimensional embeddings and multi-scale prototypes through a single-objective optimization. It integrates differentiable GW distance, distributional projection, and end-to-end learning, theoretically establishing the intrinsic equivalence between dimensionality reduction and clustering in GW space. Contribution/Results: Evaluated on multi-source image and genomic datasets, the method simultaneously enhances interpretability of dimensionality reduction and accuracy of clustering. It successfully identifies cross-scale, semantically coherent low-dimensional prototypes, demonstrating both effectiveness and generalizability of joint multi-scale structural modeling.
This work addresses the challenge of geometrically consistent matching of topological features—particularly persistent homology cycles—across disparate data systems. We propose Topological Optimal Transport (TpOT), the first framework that deeply integrates optimal transport with persistent homology. TpOT constructs a measure-topological network and defines a differentiable, geometry-aware topological-geometric joint distance within its non-negatively curved geodesic metric space, leveraging hypergraph optimal transport, measure theory, and Riemannian-geometric optimization of transport plans. On point cloud data, TpOT significantly reduces topological distortion while producing geometrically plausible and interpretable cycle-level correspondences. Theoretically, we prove that the proposed distance satisfies all metric axioms. TpOT establishes the first differentiable matching paradigm for topological data analysis that simultaneously ensures geometric fidelity and topological faithfulness.
To address distribution shift and computational scalability challenges in unsupervised domain adaptation, this paper proposes an optimal transport (OT) method formulated directly in the parameter space of Gaussian mixture models (GMMs). Unlike conventional sample-level Wasserstein OT—whose complexity is cubic in sample size—our approach is the first to model OT at the GMM parameter level, yielding a closed-form transport mapping with approximate time complexity of *O(K²d)*, where *K* is the number of mixture components and *d* the feature dimension. By operating on distribution parameters rather than raw samples, the method avoids explicit sample dependence, enabling efficient, scalable distribution alignment in high-dimensional and large-scale settings. It offers strong theoretical interpretability and integrates naturally into shallow-domain-adaptation frameworks. Extensive evaluation across 85 cross-domain tasks on nine benchmark datasets demonstrates consistent and significant improvements over state-of-the-art shallow adaptation methods.
本文使用最优传输方法(包括Wasserstein、Gromov-Wasserstein和Bures-Wasserstein距离)解决网络比较问题,并通过合成数据集和真实时间序列网络评估这些方法。
This work addresses the susceptibility of traditional model-based clustering to poor local optima arising from log-likelihood maximization. To mitigate this issue, the authors propose a novel loss function grounded in entropy-regularized optimal transport, which replaces the conventional log-likelihood objective. This new formulation preserves consistency with the global optimum while substantially reducing spurious local minima, thereby yielding a smoother optimization landscape. Within an expectation-maximization (EM) framework, they develop an efficient Sinkhorn-EM algorithm to optimize the proposed objective. Experimental results demonstrate that the method outperforms standard log-likelihood-based approaches in both C. elegans microscopy image segmentation and spatial transcriptomics clustering, achieving markedly improved clustering stability and accuracy.
This work addresses the challenge of clustering with incomplete multi-view data by proposing an end-to-end framework based on straight-path flow matching. The method completes missing views by constructing a deterministic ordinary differential equation (ODE) flow between observed and unobserved views, and enhances cross-view clustering consistency through cluster-level alignment and entropy regularization. Notably, it is the first to apply straight probability paths to incomplete multi-view clustering and provides theoretical justification that, compared to diffusion models, deterministic flows better preserve cluster structure under finite-step integration—making them more aligned with clustering objectives. The approach achieves state-of-the-art performance on standard incomplete multi-view clustering (IMVC) benchmarks.
This work addresses a critical limitation in existing optimal transport–based short text clustering methods, which often neglect semantic consistency among samples, leading to semantically similar instances being assigned disparate pseudo-labels and thereby degrading clustering performance. To overcome this issue, the paper proposes a novel clustering framework that jointly models local semantic consistency and global cluster structure for the first time. Specifically, an instance-level attention mechanism is introduced to capture neighborhood semantic relationships, which are then integrated into the optimal transport process to generate pseudo-labels that balance both local coherence and global structural integrity. Coupled with pseudo-label–guided self-supervised learning, the proposed method achieves significant performance gains over state-of-the-art approaches across multiple benchmark datasets.
本文通过将多种聚类方法统一表达为受约束的低秩投影,建立了一个通用优化框架,并提供了关于这些方法稳定性和恢复保证的理论分析。