Score
Design and implement algorithms or pipelines that automatically generate concise, human-interpretable labels or representative keywords for groups produced by clustering, using rule-based heuristics, thresholding, or other automatic annotation techniques. Build methods to select and rank candidate terms, enforce consistency and temporal stability of labels, and evaluate label quality and representativeness.
High-stakes domains—such as healthcare and finance—demand interpretable clustering outcomes to ensure transparency, accountability, and regulatory compliance. Method: This survey systematically analyzes over 120 scholarly works, proposing the first unified taxonomy of interpretability dimensions for clustering. It rigorously distinguishes intrinsically interpretable models—including rule-based, prototype-based, and sparsity-driven approaches—from post-hoc explanation techniques—such as visualization, feature attribution, and local surrogate modeling. The study further develops a use-case-oriented, structured classification framework and principled evaluation criteria. Contribution/Results: It introduces the first practical guideline for selecting appropriate interpretable clustering methods based on application requirements. The work bridges theoretical foundations with real-world deployment, providing both conceptual clarity and actionable insights to support the development and adoption of clustering algorithms that jointly optimize accuracy and interpretability—thereby advancing trustworthy AI in ethically and regulatorily sensitive contexts.
Automated label generation for scientific literature clustering faces a trade-off: traditional methods yield concise yet opaque labels, while large language models (e.g., ChatGPT) produce highly readable descriptive labels lacking theoretical grounding, systematic methodology, and empirical validation. This paper introduces the first dichotomous framework distinguishing *feature-based* from *descriptive* labels. We formalize descriptive labels, define principled evaluation metrics, and propose a structured generation pipeline integrating cluster analysis, text feature extraction, and readability optimization—fully driven by language models. Experiments demonstrate that our descriptive labels match conventional labels in interpretability and quality while bridging the theory–practice gap. The approach establishes a new paradigm for bibliometric analysis: fully automated yet human-centered, combining rigorous methodology with linguistic naturalness.
Traditional clustering methods emphasize stability of cluster centers, neglecting the practical need for stability of point-level labels—the named identifiers indicating each sample’s cluster assignment. Method: This paper formally defines “label consistency” as the pointwise label distance between successive clustering solutions, departing from the conventional center-stability paradigm. For the $k$-center and $k$-median problems, we propose the first theoretically grounded consistency-aware approximation algorithm. Leveraging combinatorial optimization and metric-space analysis, we design a dynamic label-distance modeling framework coupled with constrained optimization to jointly optimize clustering quality and label stability. Contribution/Results: Our algorithm achieves an $O(1)$-approximation ratio for both objectives and provides a tight upper bound on the label change rate. It establishes an optimal trade-off between clustering accuracy and label consistency, offering a novel, interpretable, and deployable paradigm for dynamic clustering.
High annotation costs and prolonged turnaround times plague NLP development, necessitating efficient and reliable data labeling paradigms. This paper proposes an LLM-powered Human-in-the-Loop (HITL) hybrid annotation framework that systematically integrates synthetic data generation, active learning, and human-AI collaboration, augmented with built-in mechanisms for annotation quality assessment, annotator management, and cost-benefit analysis. Unlike prior work—largely theoretical or narrowly scoped—this study introduces the first deployable, plug-and-play industrial-grade annotation methodology, bridging the critical gap between methodological research and real-world engineering practice. Empirical validation across multiple production NLP projects demonstrates that the framework consistently reduces annotation costs and cycle time by 30–50%, while maintaining label quality within required thresholds.
This study addresses the limited reliability and diagnostic capability of existing AutoClustering systems, which stem from a lack of interpretability regarding how meta-features influence the selection of clustering algorithms and hyperparameters. For the first time, the work systematically reviews the meta-features employed across 22 AutoClustering methods and organizes them into a coherent taxonomy. By integrating global interpretability through decision predicate graphs and local interpretability via SHAP values, the authors conduct a thorough analysis of feature contributions within meta-models. Their investigation uncovers structural biases and consistent patterns in current meta-learning strategies, revealing fundamental limitations of prevailing approaches. These insights not only expose critical shortcomings but also offer actionable interpretability guidelines for designing transparent and trustworthy unsupervised AutoML systems.
Selecting optimal clustering algorithms for large-scale data remains challenging, particularly under semi-supervised settings where only costly oracle queries yield partial ground-truth labels. Method: This paper introduces and formalizes the novel “scale generalization” theory: if an algorithm is optimal on a small subsample, it remains optimal on the full dataset. We derive sufficient conditions for scale generalization for single-linkage, k-means++, and smoothed Gonzalez k-center algorithms. Our framework integrates clustering stability analysis, sampling theory, and semi-supervised learning, employing subsampling-based evaluation coupled with oracle validation. Results: Empirical evaluation across multiple real-world datasets demonstrates that just 5% of the data suffices to identify the globally optimal algorithm with 100% accuracy. The core contribution is the first provably sound and empirically verifiable theory of scale generalization for clustering algorithms, enabling efficient, robust algorithm selection without full-label supervision.
This study addresses the inconsistency in human annotation caused by ambiguous category definitions in traditional content moderation. To resolve this, the authors propose an AI-driven constitutional annotation framework: large language models first assist humans in formulating structured, interpretable category “constitutions,” which then guide automated dual-axis labeling of intent and content safety. This approach shifts human effort from case-by-case judgments to high-level semantic definition. Evaluated on harassment, hate speech, and non-violent criminal conduct tasks, the method reduces cross-model annotation inconsistency by up to 57-fold compared to conventional paragraph-based rules and effectively exposes latent gaps in existing policy formulations.
This study addresses the lack of systematic guidance in parameter selection and result evaluation for unsupervised data grouping methods by proposing SmartIterator, an exploratory framework grounded in a six-stage visual analytics pipeline. The approach uniquely treats the complete sequence of groupings generated through parameter sweeps as the primary analytical object, integrating quality metrics, stability assessments, member confidence scores, and domain context to deliver method-specific, actionable workflows for tasks such as clustering and topic modeling. Implemented via the IteraScope visualization system—which features semantic color encoding, group embeddings, Sankey transition flows, violin plots, and repeated prototype detection using HDBSCAN—the framework demonstrates its efficacy across three diverse datasets: social media, regional statistics, and academic publications, enabling analysts to comprehensively interpret data structures and make informed decisions.
This work addresses the limitations of existing large language model (LLM)-based text clustering approaches, which typically rely on fixed pipelines that struggle to accommodate diverse corpus structures and cannot flexibly incorporate user-defined constraints such as desired cluster count or clustering intent. To overcome these challenges, the authors propose a dynamic multi-agent collaborative clustering framework, wherein a coordinator LLM adaptively orchestrates specialized agents—including proposers, synthesizers, reviewers, investigators, and critics—to construct controllable and context-aware text categorization systems. By replacing conventional static pipeline paradigms with a novel multi-agent dynamic decision-making mechanism, the method achieves state-of-the-art performance across seven public benchmarks, yielding up to a 32% relative improvement in Adjusted Rand Index (ARI) over the strongest LLM-based baseline.
Existing data mixing approaches are constrained by fixed semantic labels and a single granularity level, limiting their ability to flexibly explore combinatorial effects across multiple granularities. This work proposes a hierarchical, data-driven multi-granularity annotation framework that leverages learnable semantic transformations and a three-stage residual vector quantization scheme to generate up to 130,000 reusable hierarchical document codes, enabling dynamic navigation from coarse to fine granularities. Evaluated in a pretraining setting with 1B parameters and 25B tokens, the method—combined with an equal-subbucket coverage strategy—achieves an average performance gain of +0.0253 across 16 tasks at specific granularities, demonstrating a significant interaction effect between granularity selection and mixing strategy.
This work addresses the noise inherent in partial multi-label learning, where candidate label sets are contaminated with both relevant and irrelevant labels. To tackle this challenge, the authors propose a novel weakly supervised clustering approach that uniquely decomposes the cluster membership matrix into two components: a normalized Π component and an F component that preserves the binary nature of multi-label assignments. This decomposition enables the first effective integration of clustering with multi-label learning. The method employs a three-stage pipeline—prototype learning, confidence-adaptive weak supervision construction, and iterative clustering refinement—to achieve robustness against label noise. Extensive experiments on 24 benchmark datasets demonstrate that the proposed approach significantly outperforms six state-of-the-art methods across all evaluation metrics.