Multi-Domain Clustering via Measure Quantization

📅 2026-09-18
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文通过最小化概率度量如Sinkhorn散度来学习共享的聚类原型,解决了多领域数据的聚类问题,并采用小批量优化策略提高效率。
📝 Abstract
Clustering is a fundamental task in data analysis, typically addressed through centroid-based methods such as K-means. In this work, we present a general framework for multi-domain clustering via measure quantization: given samples from multiple domains, we learn a shared set of cluster prototypes by minimizing a probability metric, such as the Sinkhorn divergence or the Maximum Mean Discrepancy, between each domain's probability measure and the measure of prototypes. Data points are then assigned to clusters either via nearest centroid, or via optimal transport, a collaborative strategy that couples all samples within a domain. A mini-batch optimization strategy makes both fitting and assignment scalable, reducing memory and computational cost while preserving clustering performance. Experimental results on 5 multi-domain benchmarks spanning image, audio and sensor data show that our Sinkhorn-based method consistently outperforms classical and multi-domain clustering baselines, and that this advantage persists when scaling to hundreds of thousands of samples.
Problem

Research questions and friction points this paper is trying to address.

multi-domain clustering
measure quantization
shared cluster prototypes
Innovation

Methods, ideas, or system contributions that make the work stand out.

Measure Quantization
Multi-Domain Clustering
Sinkhorn Divergence
Mini-Batch Optimization
🔎 Similar Papers
2024-09-03International Scientific Technical Journal "Problems of Control and Informatics"Citations: 2
💼 Related Jobs
No related jobs found.