🤖 AI Summary
This study addresses critical bottlenecks in deploying LLM-based topic models for industrial applications, including high resource consumption, coarse topic granularity, and the absence of document-level distributions. To overcome these limitations, we propose SeLATM, a novel framework that introduces a segment-level topic generation strategy to transcend the constraints of traditional whole-document assignment. Furthermore, SeLATM incorporates an agent-driven feedback loop mechanism to iteratively refine topics, thereby enabling efficient and fine-grained analysis. Experimental evaluations across multiple datasets demonstrate that SeLATM significantly reduces computational overhead while maintaining topic quality and analytical performance superior to existing methods. Ultimately, this work effectively enhances data exploration efficiency, offering a practical and scalable solution for real-world topic modeling tasks.
📝 Abstract
Topic modeling is an effective technique for discovering hidden themes within documents and is widely used in text mining and data analysis across a variety of industry sectors. Recently, large language model (LLM)-based topic models have been emerged that prompt LLMs to generate topics then assign the topics to documents, producing more natural and human-readable topics than conventional topic modeling algorithms. However, the nature of topic assignment process causes certain drawbacks, such as the incapability to produce topic distributions over a document, too broad or narrow topics, and high resource consumption, which increases with the number and length of of documents being assigned topics. These issues are particularly critical for industrial applications, which require high-quality, in-depth analysis and the processing of large volumes of documents. In this context, this paper introduces a framework called SeLATM, which addresses these concerns by employing segment-level topic generation and topic refinement through agentic feedback loops. Experimental results on various datasets demonstrate that SeLATM significantly reduces the LLM resources compared to methods based on topic assignment process, while maintaining superior performance.