BERTopic AI-Powered Topic Modeling for NLP Projects

BERTopic is a modern topic modeling framework that addresses many limitations of traditional approaches. Developed by Maarten Grootendorst, it uses transformer-based embeddings (like BERT) to understand the semantic meaning of documents and clusters them based on their context rather than just word frequency.

Key Features

Transformer embeddings

Uses deep contextual language models to convert documents into dense, meaningful vector representations for better topic clustering.

c-TF-IDF for representation

Class-based TF-IDF highlights important words per topic by comparing them against the entire dataset, improving topic interpretability.

Dynamic topic reduction

Automatically merges similar topics into broader categories, simplifying results and allowing better control over the number of topics.

Visualizations

Interactive plots help explore topic similarities and key terms, improving understanding and presentation of topic model output.

How BERTopic Works

BERTopic creates interpretable and semantically meaningful topics by combining transformer-based embeddings with clustering and a custom TF-IDF strategy. Below is a breakdown of the steps involved.

Document Embedding (via BERT or similar)

  • Goal: Convert raw text into numerical representations that capture semantic meaning.
  • How: Each document is passed through a pre-trained language model (e.g., BERT, RoBERTa, Sentence-BERT).
  • Output: High-dimensional dense embeddings (e.g., 768-dimension vectors for BERT).
  • Why it matters: These embeddings preserve contextual meaning, allowing for better clustering than simple Bag-of-Words.

Dimensionality Reduction (UMAP)

  • Goal: Reduce the high-dimensional embeddings to a lower-dimensional space for clustering.
  • How: BERTopic uses UMAP (Uniform Manifold Approximation and Projection), a nonlinear dimensionality reduction technique.
  • Output: 2D or 5D representation of each document that retains semantic structure.
  • Why it matters: Lower-dimensional space improves clustering accuracy and speed.

Clustering (HDBSCAN)

  • Goal: Group similar documents into clusters that will become topics.
  • How: Uses HDBSCAN (Hierarchical Density-Based Spatial Clustering of Applications with Noise), which:
  • Automatically determines the number of clusters
  • Handles noise and outliers
  • Output: Cluster labels assigned to documents (e.g., Topic 1, Topic 2, -1 for noise)
  • Why it matters: It allows unsupervised, flexible clustering that works well for varied data distributions.

Topic Representation (c-TF-IDF)

  • Goal: Create human-readable topic labels by identifying representative words.
  • How: Applies class-based TF-IDF (c-TF-IDF):
  • Treats each cluster as a “class” or “document”
  • Calculates TF-IDF values within clusters instead of globally
  • Output: Top keywords per topic
  • Why it matters: Ensures that terms strongly associated with a specific cluster are prioritized, making topic descriptions more interpretable.

isualization (Optional)

  • Goal: Help users explore topics, their relationships, and importance.
  • How: BERTopic offers several visualizations:
  • Intertopic distance map using pyLDAvis
  • Bar charts for top terms per topic
  • Topic hierarchy (tree of merged/split topics)
  • Why it matters: Aids in understanding topic relevance and overlaps.

Installation Guide for BERTopic

System Requirements

To successfully install and run BERTopic, ensure your system meets the following basic requirements:

  • Python Version: Python 3.7 or above
  • Operating System: Works on Windows, macOS, and Linux
  • Memory (RAM): At least 4GB RAM (8GB+ recommended for large datasets)
  • Internet Access: Required to download transformer models from Hugging Face

BERTopic Optional Dependencies

Dependency Purpose Installation Command
UMAP Reduces the dimensionality of embeddings pip install umap-learn
HDBSCAN Clusters the reduced embeddings pip install hdbscan
Plotly Creates interactive topic visualizations pip install plotly
SentenceTransformers Allows use of powerful BERT-like embeddings pip install sentence-transformers
scikit-learn Core library used for ML tasks pip install scikit-learn

Advanced Usage

Custom Embeddings

By default, BERTopic uses a pre-defined transformer model internally. However, you can provide your own embeddings, typically using the sentence-transformers library, which gives more control over performance and quality.

Why use custom embeddings?

  • Better domain-specific performance (e.g., legal, medical)
  • Faster computation with smaller models
  • Ability to use multilingual or specialized models
				
					from sentence_transformers import SentenceTransformer
from bertopic import BERTopic

# Load a custom embedding model
embedding_model = SentenceTransformer('all-MiniLM-L6-v2')

# Pass it into BERTopic
topic_model = BERTopic(embedding_model=embedding_model)
topics, probs = topic_model.fit_transform(docs)
				
			

Fine-Tuning Clustering Parameters

BERTopic uses HDBSCAN for clustering and UMAP for dimensionality reduction. You can fine-tune both to improve topic quality.

Custom HDBSCAN Example:

				
					from hdbscan import HDBSCAN
from bertopic import BERTopic

# Configure HDBSCAN with specific parameters
hdbscan_model = HDBSCAN(min_cluster_size=10, metric='euclidean', cluster_selection_method='eom')

# Pass to BERTopic
topic_model = BERTopic(hdbscan_model=hdbscan_model)
topics, probs = topic_model.fit_transform(docs)
				
			

Updating Topics

After fitting the model, you can manually update topic labels or merge/split them for better interpretability.

Manual label update:

				
					topic_model.set_topic_labels({
    0: "AI and Machine Learning",
    1: "Climate Change and Environment"
})
				
			
Visualization Purpose Key Insight
Topic Similarity Map See relationships between topics Spot overlap or redundancy
Topic Hierarchy Show topic grouping structure Organize or merge related topics
Term Frequency Charts View top keywords in each topic Understand topic meaning

Frequently Asked Questions (FAQs)

BERTopic is a topic modeling technique that uses transformer-based embeddings and class-based TF-IDF (c-TF-IDF) to create interpretable topic clusters from text data.

Unlike LDA, which relies on word frequency and probabilistic models, BERTopic uses contextual embeddings from transformers and clusters documents based on semantic similarity.

  • Transformer-based embeddings (e.g., BERT, RoBERTa)
  • Dimensionality reduction (e.g., UMAP)
  • Clustering (e.g., HDBSCAN)
  • Topic representation (c-TF-IDF)

Any collection of unstructured text such as news articles, social media posts, reviews, research papers, etc.

Any model compatible with SentenceTransformers, including BERT, RoBERTa, DistilBERT, and multilingual models.

Yes, by using multilingual or language-specific transformer models like paraphrase-multilingual-MiniLM-L12-v2.

You can install it using pip:

  • bash
  • Copy
  • Edit
  • pip install bertopic

No, but a GPU can significantly speed up the embedding generation phase.

c-TF-IDF (class-based TF-IDF) is a variation of traditional TF-IDF used to extract the most representative words for each topic by treating each cluster as a “class.”

Yes, you can pass custom embedding models using SentenceTransformers.

Yes, it performs better than traditional models like LDA on short texts, thanks to transformer embeddings.

Yes, but for very large datasets, it may require significant memory and processing time due to transformer-based embeddings.

You can use built-in visualizations like intertopic distance maps, bar charts, and topic hierarchies using Plotly.

By default, it uses HDBSCAN, but you can also use other clustering methods like KMeans.

Yes, using the reduce_topics() method to merge similar topics.

Yes, use the save() and load() methods to serialize the model to disk.

The fit_transform() method returns both topic labels and probability scores.

Try tuning HDBSCAN’s min_cluster_size or reducing dimensionality more aggressively using UMAP.

Yes, you can update topic names manually using topic_model.set_topic_labels().

Absolutely. It integrates well with libraries like spaCy, Hugging Face, and Scikit-learn.

Basic preprocessing like lowercasing, removing stopwords, and punctuation is helpful. You can also pass custom preprocessing functions.

Yes, get_representative_docs() returns the most relevant documents per topic.

Use coherence scores, topic diversity metrics, or manual inspection of top words and sample documents per topic.

Check out the official documentation, GitHub repository, and community tutorials on Medium or YouTube.

No schema found.