Skip to content

perf(db): warm vector search rows into the object-store cache after open - #1160

Open
xav-db wants to merge 3 commits into
mainfrom
warm-vector-rows-on-open
Open

xav-db wants to merge 3 commits into
mainfrom
warm-vector-rows-on-open

Conversation

@xav-db

@xav-db xav-db commented Oct 1, 2026 •

Copy link
Copy Markdown
Member

Summary

After a restart, the object-store disk tier is empty, so the first vector searches download whole 4 MiB S3 parts one at a time inside the HNSW chain of dependent reads. That is about 250 parts of payloads for a 312k × 768-d index.

Vector-memory hydration only loads upper-layer rows and SimHashes. Payloads and layer-0 lists were never warmed.

This PR adds a one-shot background warm that streams each loaded Active vector generation's search rows through the object-store tier after open.

sequenceDiagram
    participant Open as open (writer or reader)
    participant Hyd as vector memory refresh
    participant Warm as object-store warm task
    participant Tier as object-store disk tier
    participant S3
    Open->>Hyd: start refresh loop
    Open->>Warm: spawn (only if tier exists and cache_puts is set)
    Hyd-->>Warm: first refresh done
    loop each Active generation, budget = half the tier
        Warm->>Tier: scan hot lane, payloads, SimHash directory (cache_blocks off)
        Tier->>S3: fetch missing 4 MiB parts once
        S3-->>Tier: parts stored on local disk
    end
    Warm-->>Open: log "vector search rows warmed..." (targets, bytes, elapsed)
Loading
  • When it runs. It runs only when there is an object-store tier with cache_puts set, which the server sets for S3 only. Local-disk and memory-only setups don't warm.
  • What it reads (VectorRows::warm_object_store_parts). For each index, three prefix scans: the hot lane [0xF0][index] (upper rows, SimHashes, layer-0 lists), payloads [0xF1][index][0x02], and the SimHash directory [0xF1][index][0x17] when the generation has one.
  • How it scans. cache_blocks = false keeps the block cache untouched, read-ahead is one part, and one fetch task per scan. Rows are discarded: SlateDB's object-store cache keeps every part it fetches, which is the point.
  • Budget. Half the tier, charged as each range's row bytes rounded up to whole parts, in target order. A target that fails partway is charged what it read and skipped, and a scope that can't be enumerated is skipped.
  • Lifecycle.
    • close() aborts the warm; part writes are atomic.
    • wait_for_startup_cache_warm() waits on a completion channel rather than taking the handle, so close() can always stop it.
  • Refactor. Target enumeration moves out of hydration into active_vector_targets, so both passes validate generations the same way.

No stored format, key, value, WAL or cache layout changes. Reads only.

Results

Full images on EC2 (us-east-2, real S3, gp3 EBS cache), with a constrained pod of 2 CPUs, a 4g memory limit and an 8 GiB disk cache. The 312k × 768-d fixture was opened fresh with an empty cache, and each template was run alone:

First query after open Before After the warm finishes (~26 s for 1.1 GB)
global vector top-50 30.8 s 714 ms
feature-filtered vector top-50 11.7 s 699 ms
family + feature vector top-50 8.0 s 603 ms
family vector k5 4.4 s 899 ms

The "before" column is v0.0.7; "after" is all three PRs on that build.

On the local embedded bench (20% fixture, 20 ms injected per GET), cold global vector top-50 went from 582 ms with the overlap PRs alone to 153 ms with the warm.

Limits

  • No readiness gate. The server doesn't wait for the warm before reporting ready, so queries in the first ~26 s still pay part of the cold cost. Gating /readyz on the warm is a possible follow-up.
  • What's not warmed: node records, entry-candidate rows (read only to recover from a deleted entry point), generations activated after open, and tenant scopes loaded later. On a reader, parts replaced by a compaction go cold again; a writer caches the SSTs it writes.
  • Budget counts row bytes in whole parts. A range is admitted only while its parts fit, so the charge never passes the budget. Shadowed versions a scan skips aren't charged; the tier's own eviction still bounds its size.

Tests

  • Contracts (production_support/vector/hydration.rs):
    • exactly the hot lane, payloads and directory of one index are read, not other row families, indexes or scopes;
    • exact byte accounting, part rounding, and scan options recorded;
    • budgets at 0, 1, half, total−1 and total, charged across targets in order;
    • a generation without a directory;
    • a target failing after its hot lane is charged and skipped, and an unreadable target.
  • Runtime test: the warm starts only for a hybrid cache with cache_puts. Waiting leaves it to close(), and close() stops an unfinished warm without waiting.
  • Public end-to-end test: a hybrid writer over a counting store. After opening with an empty cache, the first search reads fewer SST objects with the warm (20 vs 33) and returns the same rows.

Retrigger

The PR appears safe to merge, though the budget accounting and candidate-recovery coverage merit follow-up.

Findings

  1. P2 Warm can exceed its budget ▶
  2. P2 Candidate recovery remains cold ▶

Summary

The PR adds a background pass after vector-memory refresh that enumerates loaded Active vector generations and scans selected search rows into the object-store cache. It also adds a completion wait, shutdown handling, and contract and end-to-end tests.

  • The per-range rounding can exceed the configured warm budget.
  • Entry-candidate recovery reads remain outside the selected warm ranges.
Diagram
sequenceDiagram
  participant Open
  participant Refresh as Vector memory refresh
  participant Warm as Vector part warm
  participant Store as Object-store cache
  Open->>Refresh: Start initial refresh
  Open->>Warm: Spawn when cache_puts is enabled
  Refresh-->>Warm: First pass complete
  Warm->>Store: Scan hot lane, payloads, directory
  Warm-->>Open: Close completion channel
Loading

Reviews (1) · Last reviewed commit: "fix(db): keep the object-store warm with..."

xav-db added 2 commits October 1, 2026 15:46
After a restart the object-store tier is empty, so the first searches waited
for whole-part downloads one by one inside the HNSW chain of dependent reads:
about 250 parts of payloads for a 312k x 768-d index. In front of a remote
durable store (the tier caches written SSTs), a one-shot task now streams each
loaded Active generation's layer-0 rows, payloads and SimHash directory once
after the first vector memory refresh, with cache_blocks off and one part of
read-ahead, charging at most half the tier in target order. A failed target is
logged and skipped; close() stops the task and wait_for_startup_cache_warm()
waits for it.

Target enumeration moves out of hydration into active_vector_targets so both
passes validate generations the same way. Node records, generations activated
later, and parts replaced by later compactions are not warmed.
…y close()

Review follow-ups for the startup warm:
- It scans each generation's whole hot lane (upper rows, SimHashes, layer-0
  rows) rather than layer 0 alone, so rows vector memory could not hold are
  warmed too; payloads and the SimHash directory follow.
- Each range is charged its row bytes rounded up to whole parts, the least
  the tier stores for it, and a target that fails partway is charged what it
  read before the next target gets the rest of the budget.
- A scope whose generations cannot be enumerated is logged and skipped
  instead of stopping the warm for every scope.
- wait_for_startup_cache_warm watches a completion channel instead of taking
  the task, so close() can always abort an unfinished warm.

Contracts now cover the hot lane, part rounding, scan options, a generation
without a directory and a partial failure; a runtime test covers when the
warm starts and that close() stops it. The index-build wait in the public
warm test is bounded.
.scan_prefix_with_options(self.keyspace.key(prefix), .., &options)
.await;
let (read, end) = drain_rows(rows, charged, budget).await;
charged = charged.saturating_add(read.next_multiple_of(part));

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Warm can exceed its budget
The budget check uses unrounded row bytes, but this line charges a whole part afterward. With a 4 MiB part and 2 MiB of budget remaining, accepting 2 MiB of rows can leave the warm charged 2 MiB over budget. This weakens the intended half-tier limit and can evict more existing cache data than planned. Check the rounded charge before accepting a row.

Comment on lines +870 to +878
let prefixes = [
Some(VectorKey::MemoryPrefix(VectorMemoryPrefixKey::new(
index_id,
))),
Some(VectorKey::VectorPrefix(VectorItemPrefixKey::new(index_id))),
directory.then(|| {
VectorKey::SimHashDirectoryPrefix(VectorSimHashDirectoryPrefixKey::new(index_id))
}),
];

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Candidate recovery remains cold
When a search's stored entry point is stale, it scans entry candidates and reads candidate-node rows. These prefixes cover neither kind, so candidate rows in parts not fetched by the other scans can still trigger object-store downloads during the first search after warming. Consider warming those search-path rows or documenting this limit.

Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!

The warm checked each row against the budget in raw key and value bytes but
charged each range rounded up to whole parts afterwards, so the last range
could take the charge up to one part past the budget. A row is now admitted
only when the range's bytes so far, rounded up to whole parts, still fit, so
the charge never passes the budget; a debug assertion states it. A contract
walks budgets across part boundaries.

The module docs list entry-candidate rows, read only to recover from a
deleted entry point, among what the warm leaves cold.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

1 participant