Enhancing BERTopic with Intermediate Layer Representations

📅 2025-05-10
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the suboptimal performance of default top-layer embeddings in BERTopic. We systematically investigate how intermediate-layer Transformer representations affect topic quality. Specifically, we evaluate 18 intermediate-layer embedding configurations across three heterogeneous text corpora, quantitatively assessing topic coherence and diversity, and analyzing how stopword removal strategies interact dynamically with embedding layer selection. Our key contributions are: (1) first empirical demonstration that intermediate-layer embeddings consistently outperform the default top-layer embeddings; (2) discovery that stopword impact is highly layer-dependent; and (3) identification of optimal configurations that significantly surpass the BERTopic baseline—achieving up to a 12.7% improvement in topic coherence. These findings establish a reproducible, embedding-level optimization paradigm for interpretable and robust topic modeling.

Technology Category

Natural Language Processing: Interpretability, Analysis, and Evaluation of NLP ModelsMachine Learning: Representation LearningData Mining & Knowledge Management: Intelligent Query Processing

Application Category

Web Mining and Content Analysis: Topic discovery and trackingGraph Algorithms and Modeling for the Web: Graph embeddings and representation learning for Web-related graphsSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMs
📝 Abstract
BERTopic is a topic modeling algorithm that leverages transformer-based embeddings to create dense clusters, enabling the estimation of topic structures and the extraction of valuable insights from a corpus of documents. This approach allows users to efficiently process large-scale text data and gain meaningful insights into its structure. While BERTopic is a powerful tool, embedding preparation can vary, including extracting representations from intermediate model layers and applying transformations to these embeddings. In this study, we evaluate 18 different embedding representations and present findings based on experiments conducted on three diverse datasets. To assess the algorithm's performance, we report topic coherence and topic diversity metrics across all experiments. Our results demonstrate that, for each dataset, it is possible to find an embedding configuration that performs better than the default setting of BERTopic. Additionally, we investigate the influence of stop words on different embedding configurations.
Problem

Research questions and friction points this paper is trying to address.

Evaluating embedding representations for BERTopic performance improvement
Comparing topic coherence and diversity across 18 embedding variants
Investigating stop words' impact on different embedding configurations
Innovation

Methods, ideas, or system contributions that make the work stand out.

Utilizes intermediate BERT layer representations
Evaluates 18 diverse embedding configurations
Optimizes topic coherence and diversity metrics
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
D
Dominik Koterwa
Faculty of Economic Sciences, University of Warsaw
M
Maciej 'Switala
Faculty of Economic Sciences, University of Warsaw