Conversation
The band df cutoff is a fixed number while posting lists lengthen with the corpus, so the share of the band index it skips grows without any error or log line saying so. Counting skipped postings per lookup would roughly double each band lookup's index work, so the measurement runs off the query path, as a job. getBandDfCutoffCoverage(band_df_cutoff=None) reports per band and in total the band hashes, the postings (sum of df), how many of each are over the cutoff and the fractions, plus the totals at fixed reference cutoffs (50/100/200/500/1000) so a cutoff of 0 still shows what one would skip and reports taken at different times compare directly. The cutoff evaluated is the explicit one, else STORAGE_BAND_DF_CUTOFF. MongoDbStorage counts from the (band_hash, df) index alone, hinted, as a covered scan that fetches no band document. Grouping by band_hash first keeps bucketed hashes, whose df only bucket 0 carries, counted once with their total. Where df is not trusted yet (the completeness flag is unset, or a band lacks the index) the report is refused with available=false and a message naming rebuild_band_df_index, instead of counting df-less posting lists as empty; hashes whose documents carry no df at all are reported separately. MemoryStorage counts its posting lists and says that it does not apply the cutoff when matching. Exposed end to end like the other maintenance jobs: a Worker @Remote method with progress per band, GET /band_df_cutoff_coverage with an optional band_df_cutoff (a 400 unless a non-negative integer), and McritClient.requestBandDfCutoffCoverage(). The headline is logged at INFO when the job runs.
TUNING.md gains a section under the cutoff's tuning notes on how to run and read the report, with the measurement from a 7,244-sample corpus: at 200, 46.4% of band postings (51.8M of 111.8M) sit in 0.97% of band hashes (83,235 of 8,538,312). It also records why WAND/MaxScore was not built: it needs posting lists sorted by function id, which the fill-order buckets of STORAGE_BAND_BUCKET_SIZE are not. CHANGELOG gets the matching [Unreleased] entry.
_bandDfCountPipeline grouped by band_hash before totalling. Only bucket 0 of a hash carries df, and band_hashes already counted only documents with df > 0, so the per-hash $group changed none of the numbers except band_hashes_without_df. It doubled the cost (1.38 s against 0.70 s on a 430k-hash band) and its memory grew with the hashes, crossing the 100 MB $group limit near 1.5M hashes per band - hence allowDiskUse and temp files on the server. The count is now one $group with _id None over the covered (band_hash, df) index scan: documents with df > 0 are the hashes, their df the postings, with $cond sums per threshold and the max df. No allowDiskUse; the covered plan and its test are unchanged. band_hashes_without_df is dropped. The index carries no bucket field, so a df-less document cannot be told apart as a healthy bucket above 0 or as a bucket whose bucket 0 is missing without the per-hash group this removes. The corrected docstring no longer claims counting documents would count a spilled hash once per bucket. Such a hash is not counted and not served under the cutoff, and rebuild_band_df_index did not repair it: its bucketed rebuild only updated an existing bucket 0. It now upserts bucket 0 while postings survive, as the recompute after a deletion already does, and the docs name it as the repair. A band_df_cutoff above 2**63 - 1 passed the route and failed the job with OverflowError when encoded to BSON. The route answers 400 and the storage raises ValueError past BAND_DF_CUTOFF_MAX. The timing is stated as measured: a single index-only $group of exactly this shape took 39.3 s for all 20 bands on the 7,244-sample corpus (MongoDB 7.0), about 1.1 to 2 s per band. TUNING.md's overlong line in the section is rewrapped at 100.
Both the rebuild and the recompute after a pull upsert bucket 0 when it is missing while higher buckets survive. The upsert left function_ids out, and every candidate lookup reads it off each returned document, so the next match touching that hash raised KeyError and failed the job.
…t guard The band lookup treats STORAGE_BAND_DF_CUTOFF <= 0 as off, but the coverage report refused a negative setting and failed the job. It now reports it as 0. A test pins that neither the rebuild nor the recompute after a pull creates a bucket 0 for a hash without postings, and the curl example with a query string is quoted so zsh does not glob it.
This was referenced Sep 26, 2026
Since familiary#217 MemoryStorage.getCandidatesForMinHash skips a posting list longer than the job's band_df_cutoff or STORAGE_BAND_DF_CUTOFF, the same rule MongoDbStorage applies server-side. The coverage report still said this backend ignored the cutoff and appended a disclaimer to its headline, which is no longer true. Mark MemoryStorage as applying the cutoff, keep the disclaimer for a backend that does not, and test that the memory lookup drops exactly the posting lists the report counts as over the cutoff.
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Refs #201 — it doesn't answer the WAND/MaxScore question, it gives you the number to answer it with.
The worry in #201 was that a fixed
STORAGE_BAND_DF_CUTOFFskips a growing share of the index as the corpus grows (vocabulary follows Heaps' law, V(n) = 1412.8 · n^0.7247 here, while postings grow with the number of functions), and that nothing would ever say so. Recall at the corpus sizes I can reach is 1.000, so there's nothing to tune WAND against yet. What was missing was a way to watch the gap open. This adds one:It's a job, not something on the query path: counting skipped postings per lookup would roughly double the index work of every band lookup. The result gives, per band and in total, the band hashes and postings, how many of each are over the cutoff and the fractions, plus
max_df, and it repeats the over-cutoff counts and fractions at 50, 100, 200, 500 and 1000 whatever cutoff was asked about, so two reports taken months apart compare directly. On the 7,244-sample corpus:That split is the cutoff doing its job (stopword hashes are few and fat). The thing to watch is
postings_over_cutoff_fractionmoving between runs at the same cutoff; when it does, re-measure recall withbenchmarks/compare_quality.pybefore touching the cutoff.Each band is one
$groupover the(band_hash, df)index alone, a covered scan that never fetches a band document, with constant memory and noallowDiskUse. The same shape took 39.3 s for all 20 bands on that corpus (MongoDB 7.0). UnderSTORAGE_BAND_BUCKET_SIZEa spilled hash counts once with its bucket-0 total. A database whose df isn't trusted yet getsavailable: falseand a message, rather than df-less posting lists counted as empty.Two bucketing fixes came out of writing the counter:
rebuild_band_df_indexonly updated an existing bucket 0, so a hash whose bucket 0 was gone while higher buckets survived kept no df anywhere — the cutoff never served it and the report couldn't count it. The rebuild now recreates bucket 0 while postings survive, as the recompute after a deletion already did.function_idsarray, and every lookup reads it, so the next match touching that hash died withKeyError: 'function_ids'. Both upserts create it with an empty list now, and neither creates one for a hash that has no postings left.WAND/MaxScore isn't built: it needs posting lists sorted by function id, which the fill-order buckets aren't.
docs/TUNING.mdexplains how to read the report;tests/testBandDfCutoffCoverage.pychecks both backends give the same numbers, bucketing changes nothing, the count is index-only (noFETCHin the plan), and the repair cases. Unit and mongo suites,ruffandtyare green.