BERTopic AI-Powered Topic Modeling for NLP Projects
BERTopic is a modern topic modeling framework that addresses many limitations of traditional approaches. Developed by Maarten Grootendorst, it uses transformer-based embeddings (like BERT) to understand the semantic meaning of documents and clusters them based on their context rather than just word frequency.
Key Features
Transformer embeddings
Uses deep contextual language models to convert documents into dense, meaningful vector representations for better topic clustering.
c-TF-IDF for representation
Class-based TF-IDF highlights important words per topic by comparing them against the entire dataset, improving topic interpretability.
Dynamic topic reduction
Automatically merges similar topics into broader categories, simplifying results and allowing better control over the number of topics.
Visualizations
Interactive plots help explore topic similarities and key terms, improving understanding and presentation of topic model output.
How BERTopic Works
BERTopic creates interpretable and semantically meaningful topics by combining transformer-based embeddings with clustering and a custom TF-IDF strategy. Below is a breakdown of the steps involved.
Document Embedding (via BERT or similar)
- Goal: Convert raw text into numerical representations that capture semantic meaning.
- How: Each document is passed through a pre-trained language model (e.g., BERT, RoBERTa, Sentence-BERT).
- Output: High-dimensional dense embeddings (e.g., 768-dimension vectors for BERT).
- Why it matters: These embeddings preserve contextual meaning, allowing for better clustering than simple Bag-of-Words.
Dimensionality Reduction (UMAP)
- Goal: Reduce the high-dimensional embeddings to a lower-dimensional space for clustering.
- How: BERTopic uses UMAP (Uniform Manifold Approximation and Projection), a nonlinear dimensionality reduction technique.
- Output: 2D or 5D representation of each document that retains semantic structure.
- Why it matters: Lower-dimensional space improves clustering accuracy and speed.
Clustering (HDBSCAN)
- Goal: Group similar documents into clusters that will become topics.
- How: Uses HDBSCAN (Hierarchical Density-Based Spatial Clustering of Applications with Noise), which:
- Automatically determines the number of clusters
- Handles noise and outliers
- Output: Cluster labels assigned to documents (e.g., Topic 1, Topic 2, -1 for noise)
- Why it matters: It allows unsupervised, flexible clustering that works well for varied data distributions.
Topic Representation (c-TF-IDF)
- Goal: Create human-readable topic labels by identifying representative words.
- How: Applies class-based TF-IDF (c-TF-IDF):
- Treats each cluster as a “class” or “document”
- Calculates TF-IDF values within clusters instead of globally
- Output: Top keywords per topic
- Why it matters: Ensures that terms strongly associated with a specific cluster are prioritized, making topic descriptions more interpretable.
isualization (Optional)
- Goal: Help users explore topics, their relationships, and importance.
- How: BERTopic offers several visualizations:
- Intertopic distance map using pyLDAvis
- Bar charts for top terms per topic
- Topic hierarchy (tree of merged/split topics)
- Why it matters: Aids in understanding topic relevance and overlaps.
Installation Guide for BERTopic
System Requirements
To successfully install and run BERTopic, ensure your system meets the following basic requirements:
- Python Version: Python 3.7 or above
- Operating System: Works on Windows, macOS, and Linux
- Memory (RAM): At least 4GB RAM (8GB+ recommended for large datasets)
- Internet Access: Required to download transformer models from Hugging Face
BERTopic Optional Dependencies
| Dependency | Purpose | Installation Command |
|---|---|---|
| UMAP | Reduces the dimensionality of embeddings | pip install umap-learn |
| HDBSCAN | Clusters the reduced embeddings | pip install hdbscan |
| Plotly | Creates interactive topic visualizations | pip install plotly |
| SentenceTransformers | Allows use of powerful BERT-like embeddings | pip install sentence-transformers |
| scikit-learn | Core library used for ML tasks | pip install scikit-learn |
Advanced Usage
Custom Embeddings
By default, BERTopic uses a pre-defined transformer model internally. However, you can provide your own embeddings, typically using the sentence-transformers library, which gives more control over performance and quality.
Why use custom embeddings?
- Better domain-specific performance (e.g., legal, medical)
- Faster computation with smaller models
- Ability to use multilingual or specialized models
from sentence_transformers import SentenceTransformer
from bertopic import BERTopic
# Load a custom embedding model
embedding_model = SentenceTransformer('all-MiniLM-L6-v2')
# Pass it into BERTopic
topic_model = BERTopic(embedding_model=embedding_model)
topics, probs = topic_model.fit_transform(docs)
Fine-Tuning Clustering Parameters
BERTopic uses HDBSCAN for clustering and UMAP for dimensionality reduction. You can fine-tune both to improve topic quality.
Custom HDBSCAN Example:
from hdbscan import HDBSCAN
from bertopic import BERTopic
# Configure HDBSCAN with specific parameters
hdbscan_model = HDBSCAN(min_cluster_size=10, metric='euclidean', cluster_selection_method='eom')
# Pass to BERTopic
topic_model = BERTopic(hdbscan_model=hdbscan_model)
topics, probs = topic_model.fit_transform(docs)
Updating Topics
After fitting the model, you can manually update topic labels or merge/split them for better interpretability.
Manual label update:
topic_model.set_topic_labels({
0: "AI and Machine Learning",
1: "Climate Change and Environment"
})
| Visualization | Purpose | Key Insight |
|---|---|---|
| Topic Similarity Map | See relationships between topics | Spot overlap or redundancy |
| Topic Hierarchy | Show topic grouping structure | Organize or merge related topics |
| Term Frequency Charts | View top keywords in each topic | Understand topic meaning |
Frequently Asked Questions (FAQs)
What is BERTopic?
BERTopic is a topic modeling technique that uses transformer-based embeddings and class-based TF-IDF (c-TF-IDF) to create interpretable topic clusters from text data.
How does BERTopic differ from LDA?
Unlike LDA, which relies on word frequency and probabilistic models, BERTopic uses contextual embeddings from transformers and clusters documents based on semantic similarity.
What are the main components of BERTopic?
- Transformer-based embeddings (e.g., BERT, RoBERTa)
- Dimensionality reduction (e.g., UMAP)
- Clustering (e.g., HDBSCAN)
- Topic representation (c-TF-IDF)
What kind of data is suitable for BERTopic?
Any collection of unstructured text such as news articles, social media posts, reviews, research papers, etc.
Which transformer models does BERTopic support?
Any model compatible with SentenceTransformers, including BERT, RoBERTa, DistilBERT, and multilingual models.
Can BERTopic work with non-English texts?
Yes, by using multilingual or language-specific transformer models like paraphrase-multilingual-MiniLM-L12-v2.
Can BERTopic work with non-English texts?
You can install it using pip:
- bash
- Copy
- Edit
- pip install bertopic
Do I need a GPU to use BERTopic?
No, but a GPU can significantly speed up the embedding generation phase.
What is c-TF-IDF?
c-TF-IDF (class-based TF-IDF) is a variation of traditional TF-IDF used to extract the most representative words for each topic by treating each cluster as a “class.”
Can I use custom embeddings with BERTopic?
Yes, you can pass custom embedding models using SentenceTransformers.
Does BERTopic work well with short texts?
Yes, it performs better than traditional models like LDA on short texts, thanks to transformer embeddings.
Is BERTopic suitable for large datasets?
Yes, but for very large datasets, it may require significant memory and processing time due to transformer-based embeddings.
How can I visualize topics in BERTopic?
You can use built-in visualizations like intertopic distance maps, bar charts, and topic hierarchies using Plotly.
What clustering algorithm does BERTopic use?
By default, it uses HDBSCAN, but you can also use other clustering methods like KMeans.
Can I reduce the number of topics in BERTopic?
Yes, using the reduce_topics() method to merge similar topics.
Can I save and reload a BERTopic model?
Yes, use the save() and load() methods to serialize the model to disk.
How do I get topic probabilities or relevance scores?
The fit_transform() method returns both topic labels and probability scores.
Why am I getting too many small or noisy topics?
Try tuning HDBSCAN’s min_cluster_size or reducing dimensionality more aggressively using UMAP.
Is there a way to name topics manually?
Yes, you can update topic names manually using topic_model.set_topic_labels().
Can I use BERTopic in a pipeline with other NLP tools?
Absolutely. It integrates well with libraries like spaCy, Hugging Face, and Scikit-learn.
How do I preprocess data before using BERTopic?
Basic preprocessing like lowercasing, removing stopwords, and punctuation is helpful. You can also pass custom preprocessing functions.
Can I extract representative documents for each topic?
Yes, get_representative_docs() returns the most relevant documents per topic.
How do I evaluate the quality of BERTopic results?
Use coherence scores, topic diversity metrics, or manual inspection of top words and sample documents per topic.
Where can I find more learning resources about BERTopic?
Check out the official documentation, GitHub repository, and community tutorials on Medium or YouTube.