Skip to content

Report how much of the band index the df cutoff skips - #233

Open
r0ny123 wants to merge 11 commits into
familiary:mainfrom
r0ny123:feat/201-band-df-cutoff-coverage
Open

r0ny123 wants to merge 11 commits into
familiary:mainfrom
r0ny123:feat/201-band-df-cutoff-coverage

Conversation

@r0ny123

@r0ny123 r0ny123 commented Sep 26, 2026 •

Copy link
Copy Markdown
Collaborator

Refs #201 — it doesn't answer the WAND/MaxScore question, it gives you the number to answer it with.

The worry in #201 was that a fixed STORAGE_BAND_DF_CUTOFF skips a growing share of the index as the corpus grows (vocabulary follows Heaps' law, V(n) = 1412.8 · n^0.7247 here, while postings grow with the number of functions), and that nothing would ever say so. Recall at the corpus sizes I can reach is 1.000, so there's nothing to tune WAND against yet. What was missing was a way to watch the gap open. This adds one:

GET /band_df_cutoff_coverage                      # the configured cutoff
GET /band_df_cutoff_coverage?band_df_cutoff=500   # or any other
McritClient.requestBandDfCutoffCoverage()

It's a job, not something on the query path: counting skipped postings per lookup would roughly double the index work of every band lookup. The result gives, per band and in total, the band hashes and postings, how many of each are over the cutoff and the fractions, plus max_df, and it repeats the over-cutoff counts and fractions at 50, 100, 200, 500 and 1000 whatever cutoff was asked about, so two reports taken months apart compare directly. On the 7,244-sample corpus:

cutoff 200 over the cutoff total share
band hashes 83,235 8,538,312 0.97 %
postings 51.8 M 111.8 M 46.4 %

That split is the cutoff doing its job (stopword hashes are few and fat). The thing to watch is postings_over_cutoff_fraction moving between runs at the same cutoff; when it does, re-measure recall with benchmarks/compare_quality.py before touching the cutoff.

Each band is one $group over the (band_hash, df) index alone, a covered scan that never fetches a band document, with constant memory and no allowDiskUse. The same shape took 39.3 s for all 20 bands on that corpus (MongoDB 7.0). Under STORAGE_BAND_BUCKET_SIZE a spilled hash counts once with its bucket-0 total. A database whose df isn't trusted yet gets available: false and a message, rather than df-less posting lists counted as empty.

Two bucketing fixes came out of writing the counter:

  • rebuild_band_df_index only updated an existing bucket 0, so a hash whose bucket 0 was gone while higher buckets survived kept no df anywhere — the cutoff never served it and the report couldn't count it. The rebuild now recreates bucket 0 while postings survive, as the recompute after a deletion already did.
  • The recompute's upsert created bucket 0 without a function_ids array, and every lookup reads it, so the next match touching that hash died with KeyError: 'function_ids'. Both upserts create it with an empty list now, and neither creates one for a hash that has no postings left.

WAND/MaxScore isn't built: it needs posting lists sorted by function id, which the fill-order buckets aren't. docs/TUNING.md explains how to read the report; tests/testBandDfCutoffCoverage.py checks both backends give the same numbers, bucketing changes nothing, the count is index-only (no FETCH in the plan), and the repair cases. Unit and mongo suites, ruff and ty are green.

The band df cutoff is a fixed number while posting lists lengthen with
the corpus, so the share of the band index it skips grows without any
error or log line saying so. Counting skipped postings per lookup would
roughly double each band lookup's index work, so the measurement runs
off the query path, as a job.

getBandDfCutoffCoverage(band_df_cutoff=None) reports per band and in
total the band hashes, the postings (sum of df), how many of each are
over the cutoff and the fractions, plus the totals at fixed reference
cutoffs (50/100/200/500/1000) so a cutoff of 0 still shows what one
would skip and reports taken at different times compare directly. The
cutoff evaluated is the explicit one, else STORAGE_BAND_DF_CUTOFF.

MongoDbStorage counts from the (band_hash, df) index alone, hinted, as
a covered scan that fetches no band document. Grouping by band_hash
first keeps bucketed hashes, whose df only bucket 0 carries, counted
once with their total. Where df is not trusted yet (the completeness
flag is unset, or a band lacks the index) the report is refused with
available=false and a message naming rebuild_band_df_index, instead of
counting df-less posting lists as empty; hashes whose documents carry no
df at all are reported separately. MemoryStorage counts its posting
lists and says that it does not apply the cutoff when matching.

Exposed end to end like the other maintenance jobs: a Worker @Remote
method with progress per band, GET /band_df_cutoff_coverage with an
optional band_df_cutoff (a 400 unless a non-negative integer), and
McritClient.requestBandDfCutoffCoverage(). The headline is logged at
INFO when the job runs.
TUNING.md gains a section under the cutoff's tuning notes on how to run
and read the report, with the measurement from a 7,244-sample corpus: at
200, 46.4% of band postings (51.8M of 111.8M) sit in 0.97% of band hashes
(83,235 of 8,538,312). It also records why WAND/MaxScore was not built:
it needs posting lists sorted by function id, which the fill-order
buckets of STORAGE_BAND_BUCKET_SIZE are not.

CHANGELOG gets the matching [Unreleased] entry.
_bandDfCountPipeline grouped by band_hash before totalling. Only bucket 0
of a hash carries df, and band_hashes already counted only documents
with df > 0, so the per-hash $group changed none of the numbers except
band_hashes_without_df. It doubled the cost (1.38 s against 0.70 s on a
430k-hash band) and its memory grew with the hashes, crossing the 100 MB
$group limit near 1.5M hashes per band - hence allowDiskUse and temp
files on the server. The count is now one $group with _id None over the
covered (band_hash, df) index scan: documents with df > 0 are the
hashes, their df the postings, with $cond sums per threshold and the
max df. No allowDiskUse; the covered plan and its test are unchanged.

band_hashes_without_df is dropped. The index carries no bucket field, so
a df-less document cannot be told apart as a healthy bucket above 0 or
as a bucket whose bucket 0 is missing without the per-hash group this
removes. The corrected docstring no longer claims counting documents
would count a spilled hash once per bucket.

Such a hash is not counted and not served under the cutoff, and
rebuild_band_df_index did not repair it: its bucketed rebuild only
updated an existing bucket 0. It now upserts bucket 0 while postings
survive, as the recompute after a deletion already does, and the docs
name it as the repair.

A band_df_cutoff above 2**63 - 1 passed the route and failed the job
with OverflowError when encoded to BSON. The route answers 400 and the
storage raises ValueError past BAND_DF_CUTOFF_MAX.

The timing is stated as measured: a single index-only $group of exactly
this shape took 39.3 s for all 20 bands on the 7,244-sample corpus
(MongoDB 7.0), about 1.1 to 2 s per band. TUNING.md's overlong line in
the section is rewrapped at 100.
Both the rebuild and the recompute after a pull upsert bucket 0 when it is
missing while higher buckets survive. The upsert left function_ids out, and
every candidate lookup reads it off each returned document, so the next match
touching that hash raised KeyError and failed the job.
…t guard

The band lookup treats STORAGE_BAND_DF_CUTOFF <= 0 as off, but the
coverage report refused a negative setting and failed the job. It now
reports it as 0. A test pins that neither the rebuild nor the recompute
after a pull creates a bucket 0 for a hash without postings, and the
curl example with a query string is quoted so zsh does not glob it.
claude and others added 2 commits September 28, 2026 08:50
Since familiary#217 MemoryStorage.getCandidatesForMinHash skips a posting list
longer than the job's band_df_cutoff or STORAGE_BAND_DF_CUTOFF, the same
rule MongoDbStorage applies server-side. The coverage report still said
this backend ignored the cutoff and appended a disclaimer to its
headline, which is no longer true. Mark MemoryStorage as applying the
cutoff, keep the disclaimer for a backend that does not, and test that
the memory lookup drops exactly the posting lists the report counts as
over the cutoff.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

2 participants