Conversation
After a restart the object-store tier is empty, so the first searches waited for whole-part downloads one by one inside the HNSW chain of dependent reads: about 250 parts of payloads for a 312k x 768-d index. In front of a remote durable store (the tier caches written SSTs), a one-shot task now streams each loaded Active generation's layer-0 rows, payloads and SimHash directory once after the first vector memory refresh, with cache_blocks off and one part of read-ahead, charging at most half the tier in target order. A failed target is logged and skipped; close() stops the task and wait_for_startup_cache_warm() waits for it. Target enumeration moves out of hydration into active_vector_targets so both passes validate generations the same way. Node records, generations activated later, and parts replaced by later compactions are not warmed.
…y close() Review follow-ups for the startup warm: - It scans each generation's whole hot lane (upper rows, SimHashes, layer-0 rows) rather than layer 0 alone, so rows vector memory could not hold are warmed too; payloads and the SimHash directory follow. - Each range is charged its row bytes rounded up to whole parts, the least the tier stores for it, and a target that fails partway is charged what it read before the next target gets the rest of the budget. - A scope whose generations cannot be enumerated is logged and skipped instead of stopping the warm for every scope. - wait_for_startup_cache_warm watches a completion channel instead of taking the task, so close() can always abort an unfinished warm. Contracts now cover the hot lane, part rounding, scan options, a generation without a directory and a partial failure; a runtime test covers when the warm starts and that close() stops it. The index-build wait in the public warm test is bounded.
| .scan_prefix_with_options(self.keyspace.key(prefix), .., &options) | ||
| .await; | ||
| let (read, end) = drain_rows(rows, charged, budget).await; | ||
| charged = charged.saturating_add(read.next_multiple_of(part)); |
There was a problem hiding this comment.
Warm can exceed its budget
The budget check uses unrounded row bytes, but this line charges a whole part afterward. With a 4 MiB part and 2 MiB of budget remaining, accepting 2 MiB of rows can leave the warm charged 2 MiB over budget. This weakens the intended half-tier limit and can evict more existing cache data than planned. Check the rounded charge before accepting a row.
| let prefixes = [ | ||
| Some(VectorKey::MemoryPrefix(VectorMemoryPrefixKey::new( | ||
| index_id, | ||
| ))), | ||
| Some(VectorKey::VectorPrefix(VectorItemPrefixKey::new(index_id))), | ||
| directory.then(|| { | ||
| VectorKey::SimHashDirectoryPrefix(VectorSimHashDirectoryPrefixKey::new(index_id)) | ||
| }), | ||
| ]; |
There was a problem hiding this comment.
Candidate recovery remains cold
When a search's stored entry point is stale, it scans entry candidates and reads candidate-node rows. These prefixes cover neither kind, so candidate rows in parts not fetched by the other scans can still trigger object-store downloads during the first search after warming. Consider warming those search-path rows or documenting this limit.
Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!
The warm checked each row against the budget in raw key and value bytes but charged each range rounded up to whole parts afterwards, so the last range could take the charge up to one part past the budget. A row is now admitted only when the range's bytes so far, rounded up to whole parts, still fit, so the charge never passes the budget; a debug assertion states it. A contract walks budgets across part boundaries. The module docs list entry-candidate rows, read only to recover from a deleted entry point, among what the warm leaves cold.
Summary
After a restart, the object-store disk tier is empty, so the first vector searches download whole 4 MiB S3 parts one at a time inside the HNSW chain of dependent reads. That is about 250 parts of payloads for a 312k × 768-d index.
Vector-memory hydration only loads upper-layer rows and SimHashes. Payloads and layer-0 lists were never warmed.
This PR adds a one-shot background warm that streams each loaded Active vector generation's search rows through the object-store tier after open.
sequenceDiagram participant Open as open (writer or reader) participant Hyd as vector memory refresh participant Warm as object-store warm task participant Tier as object-store disk tier participant S3 Open->>Hyd: start refresh loop Open->>Warm: spawn (only if tier exists and cache_puts is set) Hyd-->>Warm: first refresh done loop each Active generation, budget = half the tier Warm->>Tier: scan hot lane, payloads, SimHash directory (cache_blocks off) Tier->>S3: fetch missing 4 MiB parts once S3-->>Tier: parts stored on local disk end Warm-->>Open: log "vector search rows warmed..." (targets, bytes, elapsed)cache_putsset, which the server sets for S3 only. Local-disk and memory-only setups don't warm.VectorRows::warm_object_store_parts). For each index, three prefix scans: the hot lane[0xF0][index](upper rows, SimHashes, layer-0 lists), payloads[0xF1][index][0x02], and the SimHash directory[0xF1][index][0x17]when the generation has one.cache_blocks = falsekeeps the block cache untouched, read-ahead is one part, and one fetch task per scan. Rows are discarded: SlateDB's object-store cache keeps every part it fetches, which is the point.close()aborts the warm; part writes are atomic.wait_for_startup_cache_warm()waits on a completion channel rather than taking the handle, soclose()can always stop it.active_vector_targets, so both passes validate generations the same way.No stored format, key, value, WAL or cache layout changes. Reads only.
Results
Full images on EC2 (us-east-2, real S3, gp3 EBS cache), with a constrained pod of 2 CPUs, a 4g memory limit and an 8 GiB disk cache. The 312k × 768-d fixture was opened fresh with an empty cache, and each template was run alone:
The "before" column is v0.0.7; "after" is all three PRs on that build.
On the local embedded bench (20% fixture, 20 ms injected per GET), cold global vector top-50 went from 582 ms with the overlap PRs alone to 153 ms with the warm.
Limits
/readyzon the warm is a possible follow-up.Tests
production_support/vector/hydration.rs):cache_puts. Waiting leaves it toclose(), andclose()stops an unfinished warm without waiting.The PR appears safe to merge, though the budget accounting and candidate-recovery coverage merit follow-up.
Findings
Summary
The PR adds a background pass after vector-memory refresh that enumerates loaded Active vector generations and scans selected search rows into the object-store cache. It also adds a completion wait, shutdown handling, and contract and end-to-end tests.
Diagram
Reviews (1) · Last reviewed commit: "fix(db): keep the object-store warm with..."