Skip to content

perf(db): overlap cold block fetches in batch point reads - #1158

Open
xav-db wants to merge 7 commits into
mainfrom
overlap-cold-batch-reads
Open

xav-db wants to merge 7 commits into
mainfrom
overlap-cold-batch-reads

Conversation

@xav-db

@xav-db xav-db commented Oct 1, 2026 •

Copy link
Copy Markdown
Member

Summary

Cold vector and search queries on an S3-backed server waited for their block reads one at a time. SlateDB resolves a multi_get one SST, and one missed block range, at a time, and HNSW and search-result batches are random keys that rarely share a block. So a cold batch of n keys paid n serial fetches: a local disk-cache read, or a whole 4 MiB object-store part on first touch.

This PR overlaps those fetches:

  • BatchReads (crates/db/src/batch_reads.rs) is a crate-level policy that replaces VectorBatchReads.
    • Single: with no SlateDB block cache attached, a batch is still one call.
    • Concurrent: keys are sorted and split into contiguous runs of 4–32 keys, with up to 16 runs per batch in flight.
    • Process-wide cap: runs beyond the first come from a shared allowance of 64, taken without waiting. A batch that gets none reads with one call.
    • Guarantees: results keep caller order, duplicates and absence. A failed run fails the batch, and a backend returning the wrong row count fails closed.
  • Callers: every vector row batch (HNSW neighbour vectors, SimHashes, layer-0 rows, restricted exact scans) and the interpreter's multi_get_raw use it. Before, HNSW batches of 32 keys or fewer were never split, and interpreter multi-gets were never split at all.
  • Search hits and point ids are checked for existence a record batch at a time through multi_get_raw, instead of one get per hit.
  • Bench: scoped_search_bench gains:
    • a probe mode: mixed templates with fresh vectors, run sequentially or as a concurrent batch, with per-probe server cgroup, network and object-store part counters;
    • index-product mode;
    • BENCH_CACHE_DISK_MB, a server-like cache split;
    • per-shape first-query object-store GET counts.
cold batch of 32 random keys, before          after
  multi_get(32 keys)                            sort keys
    block 1 ── wait                             8 runs × 4 keys, all in flight
    block 2 ── wait                               run 1: block, block, block, block
    ...                                           run 2: block, block, block, block
    block 32 ── wait                              ...
  ≈ 32 serial fetches                           ≈ 4 serial fetches

Stacked on by #1159 (expansion and projection). No stored format, key, value, WAL or cache layout changes.

Results

Embedded bench (5% fixture, local object store with 20 ms injected per GET, server-like 8 GiB hybrid cache): cold global vector top-50 went from 506 ms to 217 ms.

Warm latency is unchanged or slightly better. On a 20% fixture, three alternating rounds gave global vector p50 of 17–24 ms vs 25–29 ms before.

Tests

  • batch_reads:
    • run sizing table;
    • caller order, duplicates and absence against one call (0–1,031 keys);
    • every run reading the transaction's snapshot and staged writes across a flush;
    • failed run and short read;
    • the extra-run allowance (0, 3, 7, 15 and 40 free).
  • Vector storage: the batch-policy contract now covers small (HNSW-sized) batches.
  • Existence checks: order across record batches, wrong kind rejected before any read, and the deadline checked per item for all six materializers.

Notes

  • The policy keys on whether a SlateDB block cache is attached, as before. Every run must read one consistent view (transaction, snapshot or request view). All production callers pass one, and the test-only live-handle arms read with one call.
  • The DB production coverage baselines may need a refresh: vector exclusions are keyed by line number in search/vector/storage/mod.rs.

Retrigger

The PR is not ready to merge until index-product works for the embedded benchmark backend.

Findings

  1. P1 Embedded product indexing is missing ▶
  2. P2 Scaled fixtures produce empty probes ▶
  3. P2 Empty summaries panic ▶

Summary

The PR introduces a shared, bounded concurrent batch-read policy for cold point reads, applies it to vector and interpreter reads, and adds scoped-search benchmark modes and probes.

  • Batch results preserve caller order while overlapping cold block fetches.
  • The new benchmark has an embedded-mode setup gap and two probe-reporting issues.
Diagram
%%{init: {'theme': 'neutral'}}%%
flowchart LR
  A[Vector or interpreter batch] --> B{SlateDB block cache?}
  B -- No --> C[One multi_get]
  B -- Yes --> D[Sort keys and size runs]
  D --> E[Acquire available extra-run permits]
  E --> F[Concurrent multi_get runs]
  F --> G[Restore caller order]
Loading

Reviews (1) · Last reviewed commit: "bench(db): probe templates as an enum"

xav-db added 6 commits October 1, 2026 15:45
…scoped search

Add a probe mode that sends mixed templates with fresh vectors, groups and
items to a running server, sequentially or as a concurrent batch, with
per-probe server cgroup, network and object-store part counters. Add an
index-product mode (equality index on Item.owner), a server-like split of
one disk-cache budget (BENCH_CACHE_DISK_MB), and first-query GET counts.
SlateDB resolves a multi_get one SST, and one missed block range, at a time,
so a cold batch of n random keys waited for n fetches in turn: disk-cache
reads, or whole object-store parts on first touch. Vector batches were only
split above 32 keys, which no HNSW expansion reaches, and interpreter
multi-gets were never split.

A crate-level BatchReads policy, derived from whether a SlateDB block cache is
attached, now serves both: it sorts the keys and reads contiguous runs of 4 to
32 keys, 16 at a time, so a 32-key HNSW batch waits for 4 fetches instead of
32. Results keep caller order, duplicates and absence; a failed run fails the
batch, and a backend returning the wrong row count fails closed. Layer-0
existence and upper-vector batches now go through the policy too.
Vector, text and point-id row materialization read one stored record per hit
in turn, so a cold top-50 waited for 50 serial fetches. Records are now read a
record batch at a time through the overlapped multi-get, keeping input order,
per-hit scores and the per-item deadline check; a hit of the wrong entity kind
fails the batch before any read.
Runs beyond a batch's first now draw from one process-wide allowance of 64,
taken without waiting, so a burst of cold requests cannot multiply its
object-store part downloads (up to 4 MiB each) without bound.

BatchReads documents that every run must read one consistent view; the
test-only live-handle arms of multi_get_raw go back to one call. The
wrong-kind existence test also proves no point read happens first.
Runs were sized from the batch length before the process-wide allowance was
consulted, so a batch granted no extra runs still split into up to 16 runs
that read one after another: extra SlateDB calls with no overlap, at the
moment a node is busiest. Runs are now sized to the overlap granted, and a
batch granted none reads with one call, as BatchReads::Single does.

Tests that pin exact call counts use a private allowance; the vector storage
test, which reads through the shared one, bounds its count instead.
Comment on lines +276 to +279
"index-product" => {
fixture::create_indexes(&backend, &["item_owner"], 0, index_deadline()).await
}
"probe" => probe::run(&backend).await,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Embedded product indexing is missing

index-product creates the Item.owner index only when BENCH_HTTP_URL is set. With an embedded backend, the same mode falls through to the query path instead. That leaves the index needed for product-scoped probes uncreated, so the benchmark setup cannot be completed in embedded mode.

Comment on lines +128 to +138
let items = TOTAL_ITEMS as usize;
vectors
.into_iter()
.enumerate()
.map(|(index, vector)| {
let round = index / Template::ALL.len();
let slot = (index + round * 3 + rng.below(Template::ALL.len())) % Template::ALL.len();
Probe {
template: Template::ALL[slot],
group: rng.below(GROUPS),
item: rng.below(items),

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Scaled fixtures produce empty probes

Product probes choose from all 4,115 reference item owners, but a scaled load creates only round(TOTAL_ITEMS * scale) items. With the documented 2% fixture, most product probes therefore target owners that do not exist. Their empty-result latencies still count as successful product-search measurements, making the benchmark results misleading.

Comment on lines +293 to +296
for (template, latencies) in by_template.iter().chain([(&"all", &all)]) {
let mut sorted = latencies.clone();
sorted.sort_by(f64::total_cmp);
let at = |fraction: f64| sorted[((sorted.len() - 1) as f64 * fraction).round() as usize];

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Empty summaries panic

If a template filter selects no probes, BENCH_PROBES=0, or every request fails, all is empty. The summary then evaluates sorted.len() - 1 and panics before printing the run's error count. This makes a failed or empty benchmark harder to diagnose.

…; empty summaries

Review follow-ups for the scoped search bench:
- index-product and probe fell through to the query mode with an embedded
  backend; both now run there too.
- Product probes picked any of the 4,115 reference items, so a scaled load,
  which names only its first round(4,115 * BENCH_SCALE) items, answered most
  of them with nothing. They now pick only loaded items.
- A summary with no successful probe indexed an empty list and panicked; it
  now reports a zero count. Tests cover both.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

1 participant