Skip to content

perf(db): read expansion parents concurrently and prefetch projection records - #1159

Open
xav-db wants to merge 5 commits into
overlap-cold-batch-readsfrom
parallel-expansion-projection-prefetch
Open

xav-db wants to merge 5 commits into
overlap-cold-batch-readsfrom
parallel-expansion-projection-prefetch

Conversation

@xav-db

@xav-db xav-db commented Oct 1, 2026 •

Copy link
Copy Markdown
Member

Summary

Stacked on #1158. Two more serial cold-read chains on the search path become overlapped.

1. Projections read one record per row. A top-50 projection over search results (for example owner, $distance) waited for 50 serial cold reads. Values, selected value maps and item projections now:

record_read decides which records those are, and row_property now uses the same function for its own record read, so the two can't drift apart. Expressions stay lazy.

2. Whole-value expansion read one parent at a time. A traversal scope like group → items → attributes waited for one adjacency read per parent. Parents are now read concurrently through the request's shared index-read budget (read_children), in windows of one record batch so consumed parents are freed. Each parent keeps its point reads, because adjacency values are merge operands. Parent order, child order and first-error order are preserved, and pull-mode expansion stays lazy.

 project_items(rows)
-  for row in rows
-    resolver = new
-    get_raw(row.record)              # one cold read per row
+  for batch in rows.chunks(256)
+    resolver = new
+    resolver.prefetch(records batch rows will read)   # one overlapped read
+    for row in batch
       resolve items

 expand(rows)
-  for row in rows
-    expansion_ids(row)               # one cold read per parent
+  for window in rows.chunks(256)
+    read_children(window, budget)    # up to 16 parents in flight, in order

Results

Embedded bench, cold first query (5% fixture, 20 ms injected per GET), PR #1158 alone vs with this PR:

Shape PR1 alone + this PR
kind-B vector within top-50 332 ms 176 ms
group+kind-B vector within top-50 228 ms 97 ms
group vector within k5 148 ms 132 ms

Warm latency improves slightly too (20% fixture, group+kind-B p50: 12–14 ms before, 7–9 ms after).

Behaviour notes

  • Error precedence. A projection's read or decode errors are now batch-scoped: one from any row of a 256-row batch is returned before an earlier row's expression error, as in a filter's batches. This is documented on the resolver.
  • Cache scope. The record cache is per batch rather than per row. It is keyed by element, so values are unchanged.

Tests

  • Projections:
    • record_read matches what row_property reads for node, edge, empty and virtual-property rows and for edge endpoint paths;
    • Project, Values and ValueMap(Selected) over more than one record batch: full output in order, no per-row point reads, one read per distinct record;
    • row-local items read nothing;
    • deadlines at the prefetch and mid-batch.
  • Expansion:
    • matches a serial walk over 300 parents (more than one window) for every direction, label and output;
    • parents read concurrently up to the budget;
    • deadlines.

Retrigger

The PR appears safe to merge, though edge endpoint-property projections retain a serial cold-read path.

Summary

This PR overlaps whole-value expansion reads across parent rows and batch-prefetches stored records for selected property projections while retaining ordered results.

  • Expansion uses the request’s shared index-read budget in record-batch windows.
  • Projections reuse a record resolver across each batch. Endpoint-property projections remain outside the new prefetch path.
Diagram
%%{init: {'theme': 'neutral'}}%%
flowchart LR
  A[Projection rows] --> B[Record batch]
  B --> C[Classify property reads]
  C --> D[Prefetch current-element records]
  D --> E[Resolve rows in order]
  F[Expansion rows] --> G[Parent window]
  G --> H[Bounded concurrent reads]
  H --> I[Emit children in parent order]
Loading

Reviews (1) · Last reviewed commit: "perf(db): box the projection prefetch so..."

xav-db added 4 commits October 1, 2026 15:46
A traversal scope such as group -> items -> attributes expanded one parent at
a time, so a cold scope of a few hundred parents waited for a few hundred
adjacency reads in turn. Parents' neighbour reads now overlap through the
request's shared index-read budget, in parent order, keeping each parent's
point reads (adjacency values are merge operands) and the first error in
parent order. Pull-mode expansion stays lazy.
…d read

Property projections read each row's record with its own point read, so a
cold top-50 projection over search results waited for 50 serial fetches.
Values, selected value maps and item projections now resolve a record batch
with one shared resolver, prefetching through the overlapped multi-get the
records that per-row resolution will read. record_read, which row_property
now uses for its own record read, decides which those are, so a prefetch
never reads a record the per-row path would skip; expressions stay lazy.
…ojections

Whole-value expansion takes parents a record batch at a time and drops them
once expanded, instead of keeping every input row alive until the end.

The resolver and projection docs state the batch-scoped cache and error
precedence. Tests cover Values and selected value maps over several
batches, projection deadlines, and expansion across windows.
The batched prefetch future (an overlapped multi-get) sat inline in every
projection future, and evaluation that recurses through projections nested
them deeply enough to overflow a test thread's stack in
randomized_recursive_counts_match_the_materialized_row_oracle.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

1 participant