LLM-Assisted Topic Reduction for BERTopic on Social Media Data

📅 2025-09-18
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
BERTopic often yields redundant topics with severe semantic overlap when applied to noisy, sparse social media text. Method: We propose a lightweight LLM-augmented topic reduction method that avoids end-to-end LLM modeling; instead, it leverages LLMs to compute semantic similarity among initial BERTopic-generated topic descriptions and iteratively merges semantically proximate topics via hierarchical clustering, guided by customized prompt engineering for automated refinement. Contribution/Results: The method preserves BERTopic’s scalability while significantly improving topic diversity and coherence. Experiments on three multilingual Twitter/X datasets show an average 12.7% improvement in topic diversity (Div@5) over state-of-the-art baselines. Moreover, the approach demonstrates strong robustness and generalization across diverse LLM scales—including Llama-3-8B and Qwen2-7B—without requiring fine-tuning or architectural modification.

Technology Category

Machine Learning: Large Multimodal Models (LMMs)Natural Language Processing: SummarizationData Mining & Knowledge Management: Conversational Systems for Recommendation & Retrieval

Application Category

Semantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsUser Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendationWeb Mining and Content Analysis: Topic discovery and tracking
📝 Abstract
The BERTopic framework leverages transformer embeddings and hierarchical clustering to extract latent topics from unstructured text corpora. While effective, it often struggles with social media data, which tends to be noisy and sparse, resulting in an excessive number of overlapping topics. Recent work explored the use of large language models for end-to-end topic modelling. However, these approaches typically require significant computational overhead, limiting their scalability in big data contexts. In this work, we propose a framework that combines BERTopic for topic generation with large language models for topic reduction. The method first generates an initial set of topics and constructs a representation for each. These representations are then provided as input to the language model, which iteratively identifies and merges semantically similar topics. We evaluate the approach across three Twitter/X datasets and four different language models. Our method outperforms the baseline approach in enhancing topic diversity and, in many cases, coherence, with some sensitivity to dataset characteristics and initial parameter selection.
Problem

Research questions and friction points this paper is trying to address.

BERTopic generates too many overlapping topics from noisy social media data
End-to-end LLM topic modeling approaches require excessive computational resources
Reducing topic quantity while maintaining quality and diversity remains challenging
Innovation

Methods, ideas, or system contributions that make the work stand out.

Combines BERTopic with LLMs for topic reduction
Uses LLMs to iteratively merge similar topics
Evaluated on Twitter datasets with four LLMs
W
Wannes Janssens
Ghent University, Research Group Data Analytics, Ghent, Belgium
M
Matthias Bogaert
Ghent University, Research Group Data Analytics, Ghent, Belgium
Dirk Van den Poel
Dirk Van den Poel
Senior Full Professor (Gewoon Hoogleraar) Data Analytics, Ghent University
predictive analyticsprescriptive analyticsanalytical customer relationship managementAI