🤖 AI Summary
This study addresses online algorithms for the NP-hard correlation k-clustering problem under the random-order model. Motivated by the classical Pivot algorithm, we design an online strategy tailored to the random vertex arrival setting and employ theoretical analysis techniques specific to the random-order framework. Our primary contributions are threefold. First, we establish the inaugural Ω(log k) lower bound on the competitive ratio. Second, we derive a polylogarithmic upper bound, thereby closing a critical theoretical gap for general values of k in the online setting. Third, we achieve a polynomial-time constant-factor approximation that refines the online competitive ratio to the polylogarithmic level, effectively overcoming the limitations inherent in traditional offline approaches.
📝 Abstract
Correlation clustering has been extensively studied over the last two decades in many different computational models owing to its wide-ranging practical applications. The goal is to compute a partition of the vertex set such that the total number of disagreements, i.e. the sum of edges between clusters, and non-edges within clusters, is minimized. In this paper, we study correlation $k$-clustering, in which the total number of clusters is restricted to $k$. Correlation $k$-clustering is NP-hard, and while previous work has presented a polynomial time approximation scheme for constant $k$, there are no known results for general $k$. Our first result is a polynomial-time constant-factor approximation algorithm for correlation $k$-clustering for general $k$. The main focus of this work is in the more challenging online setting. Noting that the best competitive ratio under adversarial arrivals is known to be $\Omega(n)$, we concentrate on the well-studied random-order model, where vertices arrive in random order and on arrival of a vertex, edges to its earlier-arrived neighbors are revealed. We prove a surprising lower bound of $\Omega(\log k)$-competitiveness for any online algorithm, which can be extended to $\Omega(\log n)$ when $k = \text{poly}(n)$. Finally, the main result of this paper is a polylogarithmic upper bound on the competitive ratio for correlation $k$-clustering, using an algorithm inspired by the classic Pivot algorithm for correlation clustering.