Score
Designs, builds, and analyzes models and end-to-end pipelines that discover, label, and segment latent thematic structure in text corpora using probabilistic, matrix‑factorization, transformer and embedding-based methods (e.g., latent topic analysis, NMF, BERTopic) as well as supervised, guided, and LLM-assisted variants. Includes selecting and tuning models with coherence and coverage metrics, injecting seeds or sentiment guidance, generating interpretable topic labels (optionally via LLMs or embedding-based heuristics), and producing reproducible topic discovery and evaluation workflows.
Traditional unsupervised text analysis suffers from poor interpretability and weak semantic coherence in data-scarce domains. To address this, we propose Recursive Topic Partitioning (RTP), the first framework that deeply integrates problem-driven binary semantic tree construction with large language models (LLMs), explicitly encoding clustering logic and transforming analytical pathways into structured prompts—thereby enabling interpretable clustering and controllable generation in a closed loop. RTP synergizes recursive semantic segmentation, contextual reasoning, and prompt engineering to jointly support topic discovery and data synthesis. Experiments demonstrate that RTP-generated semantic trees significantly outperform keyword-based methods (e.g., BERTopic) in structural quality and coherence; it achieves substantial gains in few-shot classification tasks; and it enables precise, semantics-guided controllable text generation aligned with user-specified semantic features.
This work addresses the challenge of generating coherent and interpretable topics in emerging or resource-constrained domains, where traditional topic models often falter due to the absence of external knowledge. The authors propose a novel method that leverages only intrinsic word co-occurrence statistics from the corpus to directly construct the Dirichlet prior parameters for Latent Dirichlet Allocation (LDA), without altering its generative process or relying on external resources. This approach represents the first effective translation of endogenous corpus statistics into a meaningful prior distribution, significantly enhancing topic coherence and interpretability. Empirical results demonstrate that the method achieves performance on par with state-of-the-art models that depend on external knowledge, particularly excelling in data-scarce scenarios, as validated on both textual corpora and single-cell RNA-seq datasets.
This work addresses the suboptimal performance of default top-layer embeddings in BERTopic. We systematically investigate how intermediate-layer Transformer representations affect topic quality. Specifically, we evaluate 18 intermediate-layer embedding configurations across three heterogeneous text corpora, quantitatively assessing topic coherence and diversity, and analyzing how stopword removal strategies interact dynamically with embedding layer selection. Our key contributions are: (1) first empirical demonstration that intermediate-layer embeddings consistently outperform the default top-layer embeddings; (2) discovery that stopword impact is highly layer-dependent; and (3) identification of optimal configurations that significantly surpass the BERTopic baseline—achieving up to a 12.7% improvement in topic coherence. These findings establish a reproducible, embedding-level optimization paradigm for interpretable and robust topic modeling.
This work addresses the joint optimization of large language models (LLMs) and text embedding techniques to enhance efficiency and robustness in semantic matching, clustering, and information retrieval. We propose the first unified taxonomy centered on the *interaction patterns* between LLMs and embeddings—categorizing approaches into three paradigms: LLM-augmented embeddings, LLM-as-embedder, and LLM-understanding-embeddings—thereby transcending conventional task-centric taxonomies. By integrating supervised/unsupervised embedding learning, instruction tuning, prompt engineering, representation space analysis, and interpretability methods, we construct a structured knowledge graph encompassing over 100 studies. Our framework precisely delineates capability boundaries and application scopes for each paradigm, identifies persistent limitations inherited from pre-trained language models (PLMs) and novel challenges introduced by LLMs, and provides a theoretically grounded, empirically informed roadmap for future advancement.
To address incomplete topic coverage, semantic misalignment, and low inference efficiency when directly applying large language models (LLMs) to topic modeling, this paper proposes an LLM-in-the-loop collaborative framework that synergistically integrates LLMs with neural topic models (NTMs). Our core contribution is a dynamic confidence-driven topic alignment mechanism grounded in optimal transport, which complements LLMs’ strong semantic understanding with NTMs’ efficient latent-variable representation—without requiring LLM fine-tuning and maintaining compatibility with mainstream NTM architectures. The framework preserves document representation fidelity while substantially enhancing topic interpretability and coherence. Extensive experiments on multiple benchmark datasets demonstrate consistent superiority over both existing LLM-augmented and purely neural topic modeling approaches, exhibiting strong generalization and controllable computational overhead.
Existing topic models struggle to distinguish between semantic similarity and thematic relatedness among topic words. This work introduces, for the first time, the psycholinguistic dimensions of similarity and relatedness into topic model evaluation by constructing a synthetic word-pair benchmark annotated via large language models, and training a neural scoring function to quantify these distinctions. The proposed generalizable analytical framework is systematically applied across multiple corpora and model families to assess their capacity for modeling semantic structure. Results reveal significant differences in semantic preferences among distinct topic model families and demonstrate that similarity and relatedness scores effectively predict downstream task performance, offering a novel, interpretable metric for evaluating topic coherence beyond traditional approaches.
This work addresses the challenge of achieving fine-grained, high-precision topic modeling in narrow domains—particularly distinguishing semantically similar subtopics—under constraints of low cost and interpretability. The authors propose PRISM, a framework that fine-tunes a lightweight sentence encoder using sparse labels generated by a large language model (LLM) and applies threshold-based clustering to partition the embedding space into locally interpretable topic structures. PRISM employs an LLM-guided teacher–student distillation pipeline, requiring only minimal LLM queries and sparse supervision, while also investigating how active sampling strategies influence local embedding geometry. Experiments demonstrate that PRISM significantly outperforms existing local topic models across multiple corpora and even surpasses state-of-the-art large embedding models in clustering quality, offering a compelling balance of efficiency, accuracy, and deployability.
This study addresses the limitations of existing topic models in business research, where topics are often ambiguous, weakly interpretable, and lack standardization, hindering their reliability as measurement instruments. To overcome these challenges, we propose LX Topic, a novel approach that integrates large language models into neural topic modeling through a closed-loop framework. Built upon the FASTopic architecture, LX Topic enhances topic coherence via word-level semantic alignment and confidence-weighted mechanisms while preserving the original document-topic distribution, and outputs standardized document-level topic proportions. We further implement an end-to-end web-based system that unifies topic discovery, refinement, and measurement. Experiments on large-scale Amazon and Yelp review datasets demonstrate that LX Topic consistently outperforms state-of-the-art models in topic quality, clustering, and classification performance, significantly improving interpretability, stability, and measurement validity.
This study addresses the challenge of reliably capturing narrative themes and their dynamic evolution in small-scale poetic corpora, where traditional topic models often yield unstable results. To overcome this limitation, the authors propose a lightweight hybrid framework that integrates unsupervised Latent Dirichlet Allocation (LDA) with supervised sparse Partial Least Squares Discriminant Analysis (sPLS-DA). By incorporating multi-seed consensus strategies and narrative hub analysis—and deliberately filtering out prosodic and other surface-level linguistic features—the approach enables a computationally rigorous close reading of *Eugene Onegin*. The method substantially enhances topic stability and literary interpretability within limited corpora, successfully identifying five coherent themes that align meaningfully with the poem’s emotional trajectory and narrative arc. This work thus establishes a transparent, reproducible paradigm for computational analysis of densely layered literary texts.
This work addresses the limitations of traditional neural topic models, which rely on the bag-of-words assumption, ignore contextual semantics, and are vulnerable to data sparsity. The authors propose a novel approach that leverages large language models to generate next-word probability distributions under tailored prompts, projects these distributions onto a predefined vocabulary to construct semantically rich soft labels, and uses them as supervision signals to guide the topic model in reconstructing documents from the language model’s hidden states. This is the first method to integrate language model–driven semantic soft labels into topic modeling, substantially improving topic coherence and purity. Extensive experiments demonstrate superior performance across three benchmark datasets and significant gains over existing approaches in semantic document retrieval tasks.