Feature request
_bertopic.py contains several near-identical copy-paste patterns that have diverged slightly over time. Extracting them into shared helpers would reduce duplication and prevent future divergence.
Three patterns are affected:
-
Document aggregation — documents.groupby(["Topic"], as_index=False).agg({"Document": " ".join}) appears in 4 methods: _extract_topics, hierarchical_topics, update_topics, and partial_fit. When a new column needs aggregating (e.g., images), every call site must be updated independently.
-
Feature names extraction — a 5-line sklearn version check (get_feature_names vs get_feature_names_out) is duplicated in hierarchical_topics and _c_tf_idf.
-
Topic name from words — "_".join([x[0] for x in words][:N]) appears in 4 places across 2 methods (topic_labels_ with [:4], hierarchical_topics with [:5] in 3 places), with inconsistent slicing.
Motivation
Reduce maintenance burden. Any change to the aggregation pattern (e.g., adding image support) currently requires updating 4 call sites in lock-step. The duplicated sklearn version check is a maintenance liability — when the old API is eventually dropped, two separate locations need updating. Inconsistent topic name slicing could produce subtly different results depending on the code path.
Your contribution
I can submit a PR that extracts three private helpers:
_aggregate_documents(documents, columns=None) — consolidates the groupby pattern
_get_feature_names(vectorizer) — wraps the sklearn version check
_topic_name_from_words(words, n=5) — consistent slicing
Zero behavior change. Each helper is a direct extraction of existing code. All existing tests pass without modification.
I've already been prototyping this in my fork, so I can turn it into a PR quickly if the direction looks good to you.
Feature request
_bertopic.pycontains several near-identical copy-paste patterns that have diverged slightly over time. Extracting them into shared helpers would reduce duplication and prevent future divergence.Three patterns are affected:
Document aggregation —
documents.groupby(["Topic"], as_index=False).agg({"Document": " ".join})appears in 4 methods:_extract_topics,hierarchical_topics,update_topics, andpartial_fit. When a new column needs aggregating (e.g., images), every call site must be updated independently.Feature names extraction — a 5-line sklearn version check (
get_feature_namesvsget_feature_names_out) is duplicated inhierarchical_topicsand_c_tf_idf.Topic name from words —
"_".join([x[0] for x in words][:N])appears in 4 places across 2 methods (topic_labels_with[:4],hierarchical_topicswith[:5]in 3 places), with inconsistent slicing.Motivation
Reduce maintenance burden. Any change to the aggregation pattern (e.g., adding image support) currently requires updating 4 call sites in lock-step. The duplicated sklearn version check is a maintenance liability — when the old API is eventually dropped, two separate locations need updating. Inconsistent topic name slicing could produce subtly different results depending on the code path.
Your contribution
I can submit a PR that extracts three private helpers:
_aggregate_documents(documents, columns=None)— consolidates the groupby pattern_get_feature_names(vectorizer)— wraps the sklearn version check_topic_name_from_words(words, n=5)— consistent slicingZero behavior change. Each helper is a direct extraction of existing code. All existing tests pass without modification.
I've already been prototyping this in my fork, so I can turn it into a PR quickly if the direction looks good to you.