Score
Designs, builds, and evaluates methods for partitioning sets of items into coherent groups and the search procedures that discover those partitions. This includes specifying similarity or distance measures, objective functions and constraints, algorithmic architectures (e.g., hierarchical, flat, density- or model-based), search and optimization strategies or heuristics to find groupings, and validation metrics and analyses to compare and refine grouping outcomes.
This study addresses the need for efficient enumeration of set partitions in combinatorial optimization and related domains. The authors systematically evaluate the performance of several existing enumeration algorithms and propose approximation formulas for estimating the number of set partitions that balance accuracy and computational efficiency across both small- and large-scale instances. Through comprehensive benchmarking, the algorithm by Djokic et al. is identified as the most effective in practical applications. The primary contributions include an empirical comparison of prominent set partition enumeration methods and the derivation of scalable approximation formulas that provide practitioners with clear guidance for algorithm selection and rapid estimation of partition counts.
In bi-objective combinatorial optimization, the two-phase Pareto optimization framework suffers from computational redundancy in its second phase, where independent invocations of ranking algorithms repeatedly evaluate identical solutions. To address this, we propose a coverage-driven region grouping mechanism: (i) an implicit grouping strategy modeled on solution coverage relations, eliminating explicit enumeration and redundant evaluations; and (ii) a multi-scale explicit region merging method that drastically reduces ranking algorithm calls. Integrated into a two-phase Pareto optimization framework augmented with binary search and region clustering, our approach is validated on the bi-objective minimum spanning tree problem. Experimental results show substantial reduction in second-phase solving time and significant overall efficiency gains. The core contribution lies in pioneering “coverage” as a grouping criterion—enabling synergistic implicit deduplication and structured search—thereby advancing both theoretical rigor and practical scalability in bi-objective optimization.
This work addresses the biclustering problem with pairwise constraints (must-link/cannot-link) on weighted bipartite graphs, aiming to extract *k* disjoint dense bicliques satisfying prior domain knowledge to enhance interpretability and clustering quality. Methodologically, we propose a custom branch-and-cut algorithm based on low-dimensional semidefinite programming (SDP) relaxation, complemented by an efficient heuristic integrating low-rank matrix decomposition and the augmented Lagrangian method. To boost scalability and efficiency, we introduce valid inequalities, cutting-plane strengthening, and block-coordinate projected gradient optimization. Experimental results demonstrate that our exact algorithm significantly outperforms general-purpose integer programming solvers, while the heuristic rapidly delivers high-quality solutions on large-scale instances. To the best of our knowledge, this is the first work to systematically incorporate pairwise constraints into a biclustering optimization framework—achieving both theoretical rigor and practical applicability.
This study investigates the robustness of subset rankings under ordinal aggregation by merging similar items in an item similarity graph, assuming additive evaluation metrics. The problem is formulated as four classes of combinatorial optimization tasks, aiming to maximize or minimize either the absolute or relative rank of a given subset. The work provides the first systematic characterization of the computational complexity of ranking optimization with partitioning operations, establishing NP-hardness for most variants while developing exact and approximation algorithms tailored to realistic, structured graph topologies. The proposed methodology is successfully applied to assess the robustness of rankings of greenhouse gas emission sources, demonstrating its practical utility across domains.
Subspace clustering in high-dimensional data often yields multiple semantically distinct subspaces, yet existing methods require manual specification of both the number of subspaces and the number of clusters within each—rendering them parameter-sensitive and poorly interpretable. This paper proposes an automatic, non-redundant multi-subspace clustering framework. First, it introduces the Minimum Description Length (MDL) principle to non-redundant clustering, enabling joint, adaptive inference of both the optimal number of subspaces and the cluster count per subspace. Second, it designs a split-merge-based greedy search strategy coupled with a subspace-level outlier encoding mechanism, allowing simultaneous outlier detection. Evaluated on multiple benchmark datasets, the method achieves competitive accuracy against state-of-the-art approaches while significantly improving parameter robustness, model interpretability, and practical applicability.
This work addresses the lack of a unified optimization foundation and quality evaluation mechanism in hierarchical clustering methods based on minimum-distance merging. It proposes a class of hierarchical agglomerative clustering algorithms derived from a bipartite objective function. By reinterpreting classical hierarchical clustering procedures as optimization processes of this objective, the study establishes, for the first time, a general connection between hierarchical clustering and an explicit optimization goal. This framework not only provides a unified theoretical interpretation for several existing algorithms but also naturally yields cluster quality metrics and stopping criteria, thereby enhancing both the interpretability and practical utility of hierarchical clustering.
This study addresses the feasibility of leader-induced follower responses in bilevel matching optimization, where the leader specifies mandatory and forbidden edges. Focusing on maximum-weight matchings and minimum-weight perfect matchings, the work combines parameterized complexity analysis, combinatorial optimization theory, and bilevel modeling to show that even with a single mandatory or forbidden edge, the follower response problem for general matchings remains NP-hard. In contrast, when every edge is either mandatory or forbidden, the perfect matching variant becomes polynomial-time solvable. Furthermore, the problem is shown to be fixed-parameter tractable with respect to the number of non-mandatory edges, thereby precisely delineating the complexity boundary between these two matching settings under structural constraints.
This study addresses the challenge of excessive model size and low solution efficiency in the clique partitioning problem caused by redundant transitivity constraints. The authors identify a class of globally redundant transitivity constraints within 0–1 integer linear programming formulations: although each individual constraint defines a facet of the feasible polyhedron, their collective removal—under correlation clustering instances with edge weights restricted to {−1, 1}—does not alter the set of optimal solutions. Drawing on polyhedral theory and computational experiments, the proposed simplified model substantially reduces problem scale and demonstrates markedly improved computational efficiency compared to existing modeling approaches for correlation clustering tasks.
This work addresses the node clustering problem in edge-colored hypergraphs, where the goal is to assign colors to nodes so as to maximize agreement with the colors of their incident hyperedges—a problem known to be NP-hard. The paper proposes a novel purely combinatorial approximation algorithm that, for the first time, achieves an approximation factor strictly below 2 without relying on linear programming, thereby breaking through the performance barrier of existing combinatorial approaches. By integrating local search with greedy strategies, the method significantly enhances approximation quality while preserving scalability, offering an efficient and theoretically superior solution to this challenging optimization problem.