Skip to content

feat(kv_index): the chain index, a lossless KV-event relay with state snapshots, and liveness-aware cache routing - #2814

Draft
slin1237 wants to merge 1 commit into
mainfrom
perf/kv-router-leap-pr
Draft

slin1237 wants to merge 1 commit into
mainfrom
perf/kv-router-leap-pr

Conversation

@slin1237

@slin1237 slin1237 commented Oct 6, 2026 •

Copy link
Copy Markdown
Member

One squashed commit: the loop branch's tree without its lab notes, which stay on the loop branch with the full history and the scoreboard. Heads, base and corpus are in the Provenance table at the end.

What

The gateway's KV-aware routing gets a new index, a lossless event relay and liveness-aware routing, all exact against a reference.

  • Chain index. The engines' block-hash chains are stored run-length compressed in an arena, one content hash per position shared by every holder.
  • Each run carries one coverage bit per worker plus a table of partial holders, so a popular prefix stays one run however many workers hold parts of it.
  • A chain that diverges inside a run becomes a child keyed by offset and next hash, so a request's decode tail never splits the run.
  • Readers take no locks and write to no memory: a run's window is read under a seqlock version and checked again after the read.
  • Writers lock one run at a time, descend without locks, and recycle dead runs, hash arrays and child tables through lock-free free lists.
  • Each event lane keeps an open-addressing block map, and a lane pool schedules whole workers so events of one worker stay in order.
  • A sharded form keeps one index per memory node, and a lookup unions the shards' answers exactly because worker sets are disjoint.
  • Every indexer must equal a single-threaded reference indexer on content and on every lookup score, after any replay.
  • The backend is chosen with --kv-index {positional,chain}, the positional indexer stays the default until its soak, and run is a deprecated alias of chain.
  • Relay. Both engines' wire layouts decode into one model, and a field that cannot be read costs that event, not the batch.
  • Each stream is normalized: tiers, cache groups, locality, ownership, namespaces, bigram pages, with one counter per drop reason.
  • Stores and removals are forwarded one for one, and the gateway counts physical copies per worker, rank and tier.
  • Every data-parallel rank is subscribed, and the engines' chain hashes are reproduced so a misconfigured worker shows as a mismatch rate.
  • A bounded history of 10,000 batches within 256 MiB serves resumes, and a dropped batch is refilled from the engine's replay socket.
  • The relay keeps the engine's live blocks and serves them as a state snapshot to a subscriber whose cursor predates the history.
  • The relay starts at servicer boot, primes itself from the engine's replay, and reads a publisher restart by three rules that hold on every wire.
  • The engine's load rides on every batch, with heartbeats while the engine is quiet, so load-aware routing no longer waits for the 10 s poll.
  • Hybrid models get a KV-cache group policy: sliding-window groups are dropped beside a main-attention group, window stores are read tail-aligned otherwise.
  • The vLLM servicers report queued uncached token-work, generation throughput and hit rate from their own bookkeeping.
  • Routing and recovery. One admission cursor per worker and rank handles gaps, duplicates, publisher restarts and snapshot resyncs.
  • Keepalive pings every second and a liveness tracker veto a dead, partitioned or wedged worker within about two seconds, with --worker-stall-secs 2 and --worker-wedge-secs 3.
  • A returned worker is re-admitted on first contact with a closed circuit breaker.
  • A thin or returned worker receives a slice of cache misses and a bounded share of hits.
  • Every request's end reaches the policy that placed it through the worker's load guard, on every path.
  • Worker selection runs through a cost-function selection layer whose default reproduces the existing decision.
  • Worker overload protection is on by default as steering: token usage 0.8, waiting requests 8, the least-loaded worker when every worker is over, shedding opt-in.
  • Ties within one microsecond of expected wait are drawn uniformly, after a 128-worker soak showed only 47 workers ever receiving a request.
  • The mock worker is a vLLM-style engine with the engines' event wires, fault hooks and engine truth, and its replay binary scores routing against the fleet's oracle.
  • The stdout log sink no longer blocks runtime threads: a 2.7 s TTFT p50 at info level became 115 ms.
  • jemalloc purges on schedule: idle RSS 319 MB against 438 to 570 MB before.

Why

The index was the bottleneck, the relay lost events, and failures were noticed late.

  • The positional indexer keyed one entry per block position and probed once per request block.
  • Its jump search over-counted on divergent prompts and under evicted middle blocks: 1,538 of 3,280 lookups disagreed with the reference.
  • Made exact, it cost 25 to 27 µs per lookup and 334 to 478 bytes per resident block, plus a locked hash-map write per block.
  • The old relay read six fields of one legacy layout, subscribed to rank 0 only, and lost a whole batch on any undecodable event.
  • A HiCache demote deleted a block the engine still served from host, and a fresh gateway could not learn what warm engines held.
  • Worker failures reached routing through the health checker, tens of seconds after the transport knew.

Benchmark publication

On the same cores of one host, the chain index sustains 1.44× the comparison indexer's block operations per second under the published memory setting and 1.50× with local memory, at a third of its lookup p99.

  • Every row uses the comparison router's benchmark definition and its Mooncake trace, whose parameters are in the Provenance table.
  • Sustained means fresh-process trials achieve at least 99% of the offered rate with a valid generator, capacity means the achieved rate when overloaded.
  • Every row carries the lookup p50 and p99 at that load, because an overloaded headline alone cannot see a slower read path.
  • The 1.61× and 1.56× figures reported during the loop paired the comparison's loaded-host points with ours and are withdrawn.

Same binary, equal cores, scaled layout

The chain index sustains 1.37× and reaches 1.53× the capacity of the comparison indexer on 52 lane cores, with p50 2.7× and p99 3.5× lower.

Indexer Sustained (M block ops/s) Per lane core (M) Lookup p50 / p99 at that load (µs) Capacity (M block ops/s) Lookup p50 / p99 overloaded (µs)
comparison CRTC 479.3 [478.9, 479.5] 9.22 3.5 [3.5, 3.5] / 14 [14, 14] 764.5 [734.8, 811.3] (300 ms window) 2.7 / 12
SMG chain index 654.6 [654.4, 654.6] 12.59 1.3 [1.2, 1.3] / 4 [4, 4] 1,171.0 [1,140.1, 1,227.9] (200 ms window) 1.2 / 4
SMG positional (the replaced design) 131.6 [131.4, 131.8] 2.53 27.3 [25.5, 32.2] / 413 [343, 478] 144.7 [143.9, 145.1] 23.4 / 565

Published memory setting, series medians with 95% intervals, the sustained column is the headline and capacity a within-harness figure.

  • Brackets behind the sustained points: the comparison keeps up at 480.8 M and fails at 511.2 M.
  • The chain index keeps up at 656.4 M and fails at 718.7 M, the positional at 132.4 and 143.1 M.
  • The control pairs put the noise floor at one unit in the last digit for sustained throughput and lookup p50.

Both systems at their tops of stack on a quiet host, comparison layout

These are the rows the ratios cite, series median against series median: 1.44× interleaved and 1.50× with local memory, with 1.43× and 1.50× on the kept-up points.

Memory Indexer Keeps up at (bracket) Fails at 20-trial series: achieved median [95% CI] (M) Kept up / discarded Per lane core (M) Lookup p50 / p99 (µs)
--interleave=all (published method) comparison CRTC 744.5 M (430 ms window) 800.0 M (98.9 / 99.2 / 99.3%) 739.1 [736.8, 740.8] (control 739.1 [735.6, 741.1]) 13 of 20 / 6 12.5 3.3 [3.1, 3.4] / 13.8 [13.0, 14.0]
--interleave=all SMG chain index 1,067.0 M (300 ms window) 1,163.6 M 1,063.6 [1,063.3, 1,063.9] (control 1,063.4 [1,062.8, 1,063.6]) 19 of 20 / 0 18.0 1.3 [1.2, 1.3] / 3.7 [3.4, 4.0]
local (--cpunodebind=0 --membind=0) comparison CRTC 924.0 M (346 ms window) 992.9 M (97.4 / 96.6 / 98.0%) 918.1 [916.1, 919.4] (control 917.9 [914.5, 918.9]) 16 of 22 / 3 15.5 2.7 [2.7, 2.7] / 9.9 [9.9, 10.0]
local SMG chain index 1,383.8 M (231 ms window) 1,509.0 M 1,379.3 [1,377.1, 1,380.0] (control 1,379.0 [1,377.2, 1,380.0]) 16 of 20 / 5 23.4 1.3 [1.3, 1.3] / 3.6 [3.6, 3.6]

Read the series column for the ratio and the lookup column for the latency gap, 59 lane cores for every system.

  • The chain index's kept-up points did not move with host load, while the comparison's rose from 682 M and 860 M on the loaded host.
  • Local memory is worth +24% to the comparison's kept-up point and +30% to the chain index's.
  • At the 150 ms window, 2.13 B offered, the comparison harness's generator is invalid in every trial, which is the harness's ceiling.

Lookup latency across the load ladder

At every window both systems sustain, the chain index answers in about 1.3 µs at p50 and about 4 µs at p99 against 3.3 to 3.5 and 12.6 to 14.1 µs.

Offered (window) chain index p50 / p99 (µs) comparison p50 / p99 (µs) positional p50 / p99 (µs)
107 M (3 s) 1.28 / 4.06 3.36 / 14.02 25.9 / 234
213 M (1.5 s) 1.22 / 4.10 3.36 / 13.63 not kept up (155 M)
427 M (750 ms) 1.31 / 4.00 3.52 / 14.05 not kept up (157 M), 26.6 / 508
640 M (500 ms) 1.25 / 4.13 3.39 / 13.70 not kept up (161 M)
1,067 M (300 ms) 1.31 / 4.19, kept up 3 of 3 (1,061.8 M) 2.94 / 12.58, not kept up (863.0 M) not kept up (158-165 M)

Same binary, comparison layout, measurement cores, 3 interleaved trials per window, medians.

  • On the mock fleet under soak traffic the lookup reads p50 3.13, p99 10.1 and p999 29.4 µs over 28,401 lookups.

Memory per block

The chain index holds 1/14.8 of the comparison indexer's bytes per resident block in the benchmark's shape.

Indexer Bytes per resident block Live bytes
comparison CRTC 887 (913 in the earlier row, 7% of it harness channel retention) 1,915 MB
SMG chain index, final shape 59.9 self-reported per membership (index 25.9 + lane maps 34.0), 62.5 as allocated 136 MB at the 65 B pool-head row
SMG positional 333.9 700 MB

Counting allocator in the shared harness binary at the 3 s window, 2,096,883 resident blocks for every backend.

  • The benchmark's shape has 128 holders per chain, where the shared hash arrays pay off most, the gateway's one to four.
  • There the RSS delta is 112 MB against the positional's 138 MB, 128 against 158 bytes per membership including about 25 MB of request data.
  • An 8-byte key-less lane-map slot measured 51.9 B per block and was reverted because it cost 30 to 40% of lane CPU.

Exactness

All three indexers agree with the reference on every score, count and membership, and every defect found on the way was fixed and pinned by a test.

  • The recorded streams, vLLM 0.31.0, SGLang 0.5.21, two data-parallel ranks and HiCache write-through, replay into both backends and the reference, compared after every batch.
  • Every chain is replayed as a request in full, as its first half, with its middle block replaced, and with its tail replaced and extended.
  • Seeded corpora of 300,000 events under two seeds and two shardings add evicted blocks, heals, divergent siblings, duplicate and host copies, clears and removals.
  • Sixteen concurrent event lanes with four readers and worker replacement replay into the reference at the end, 20 of 20 release runs.
  • Two randomized corpora joined: six seeds of 3,000 wild events checked against the reference every 16 events, and twins with two engine hashes per position.
  • The positional indexer before this series also disagreed on 562 of 3,497 lookups on the hole corpus.
  • Through the same reference the comparison indexer is exact without holes and under-counts 12 of 21,151 lookups under holes by 1 to 28 blocks.
  • A store moving a held engine hash left its old membership in place in both SMG indexers, now the latest store wins in all three.
  • A lane-map write race under-credited a worker after a concurrent split, failing 20 of 400 two-shard harness runs and 3 of 400 at one shard.
  • With the rule changed to reachability through the forwarding chain it runs 0 failures of 400 at either shard count, the unfixed index 44%.
  • The hot path is unchanged: one-shard lookup p50 0.58 against 0.58 µs, 52.9 against 52.8 ns of lane CPU per block.
  • A cut with nothing beyond it, 24 positions in a window-only feed, panicked a lookup through a slab sentinel, and the run now ends there.
  • The recorded sliding-window feeds, read as the engine writes them, are exact: 8,738 / 8,738, 441 / 441 and 8,745 / 8,745 blocks.
  • Their 296,314 and 12,708 lookups are identical, and the engine-hash conflicts fell from 6,267 to 131 once the head-aligned offline reading went.
  • On gpt-oss-20b with sliding-window layers the router's credit equals the engine's reuse, 0.518 and 0.503 over two phases, agreeing on 524 of 525 prompts.

The two backends in the gateway

In the gateway the chain index matches the positional path's hit rate with the lookup p99 fifty times lower and less memory.

chain index (--kv-index chain) positional (default)
hit / oracle 0.996 [0.995, 0.996] 0.994 [0.990, 0.997]
in-gateway lookup p50 / p99 (µs) 3.1 / 7.9 25.5 / 438
event apply per batch p50 / p99 (µs) 18.8 / 230 39 / 722
gateway RSS peak / final (MB) 676 / 443 711 / 598
TTFT p50 / p99 (ms) 181 / 3,430 178 / 5,840

Mock fleet of 8 realistic workers, Mooncake rows 0-3,999 at 3×, cache_aware, 3 interleaved runs each, medians, 0 errors, goodput and reuse inside the unlocked noise.

  • Soak s6 at 8 workers served 7 of 7 windows at hit/oracle 0.994, per-minute minimum 0.952, 301,419 batches applied with 0 stale.
  • Its idle memory was 99 MB allocated and 242 MB RSS, against the positional soak's 212 MB allocated.
  • Soak s7 at 128 workers served 94 of 94 windows over twelve hours, 54 fault cycles, 6.74 M KV batches, 0 subscription failures, 12 preemptions.
  • Its hit/oracle fell from 0.999 to 0.966 by hour twelve, each restart on a routed worker dropping its memberships for good until the refill fix.
  • Idle memory settled from 213 to 202 to 204 MB allocated, RSS 366 to 378 MB flat over twelve hours, 2.9 MB per worker.
  • That is about 92 B per membership all-in over 1,507,342 memberships, s6 holding 262k.
  • Soak s8 on the fixed head at six hours served 47 of 47 windows at hit/oracle mean 0.996, per-minute minimum 0.945, no alerts.
  • Idle RSS by hour was 245 to 254 MB, median TTFT p50 / p99 144 / 1,352 ms, below s6's first six hours.
  • The index side of the flip evidence is clean, and the one routing defect it showed, an emptied worker left unfilled, is fixed here.
  • Against s6 hour by hour s8 holds hit/oracle 0.996 to 0.982, goodput 10.6 to 11.3 against 7.6 to 10.4 req/s, 0 preemptions against 284.

s6 showed the lookup cost growing with churn at constant memberships, and children at any offset fixed it in this series.

Churn evidence Under churn before the fix After children at any offset
churn bench, runs live 6k → 859k at unchanged content, mean run length 61 → 2.9 blocks 550k, branch splits 482,878 → 0
churn bench, runs walked per lookup 5.4 → 60.8 2.85
churn bench, lookup p50 / p99 (µs) 1.1 / 4.5 → 25.06 / 127.3 1.41 / 5.4, exact at every sample
soak s6, 8 workers, lookup p50 / p99 (µs) hour 1 3.12 / 9.99, hour 2 4.96 / 31.9, 7.99 / 58.3 and p999 128 after a traffic pause, plateau 7.5 / 36 to 40 from hour 3 run c1, two hours on 8 workers, 3.0 to 3.3 / 7.8 to 9.9 flat, 970 to 1,214 runs live for 234 to 274k blocks at hit/oracle 0.989 to 1.000, then c3 3.6 to 5.0 / 8.0 to 13.9 flat at 0.978 to 1.000, s8 4.4 to 4.7 / 12.5 to 14.3 flat over six hours, c4 on the tuned head 3.7 to 3.9 / 15.2 to 15.6 for forty minutes then 5.0 to 5.7 / 26.7 to 29.1 beside a gate's builds
soak s7, 128 workers, lookup p50 / p99 (µs) 4.6 / 15.2 to 5.2 / 15.7 over twelve hours, p99 flat

The churn bench stores decode blocks under the prompt's last block, as engines do.

Fault drills

Every drill of the recovery protocol passes in both topologies, and a router restart on real engines is a 0.6 s outage with the index rebuilt in under two seconds.

Drill Topology Pass rule Measured
engine killed direct out of routing ≤ 2.5 s, no hung stream out in 2.1 s, slowest request 0.70 s
worker partitioned direct out ≤ 2.5 s, back after the heal out 2.0 s, back 0.1 s
worker unreachable direct out ≤ 2.5 s, back after the listener returns out 2.2 s, back 0.3 s
engine restarted, empty cache direct out ≤ 2.5 s, re-admitted and resubscribed ≤ 2.5 s, first request ≤ 10 s into a trickle out 2.0 s, first request 0.3 s into the trickle
engine paused 8 s, health answering direct vetoed within the pause, routable after resume wedged 3.1 s into the pause, routable 0.0 s after resume
gateway restarted under load direct serving ≤ 10 s, p99 ≤ 2× steady + 0.5 s serving in 0.8 s, p99 0.58 → 0.64 s
20 batches dropped on the wire direct gap detected and replayed gap seen 0.6 s after the drop, 20 recovered
publisher delayed 1.5 s direct mean apply lag ≥ 0.75 s 1.49 s over 267 batches
publisher restarted, cache kept, a hits-only phase then the trickle direct resync counted, the emptied worker routed to again under hits-only traffic and under the trickle resync 0.2 s, routed again 1.0 s into the hits-only phase and 0.1 s into the trickle
engine frozen 8 s (SIGSTOP) direct back after SIGCONT back 0.0 s, slowest 2.05 s
engine CPU-starved 15 s direct healthy after the hogs healthy 0.0 s after, slowest 0.70 s
overload, 512 streams direct no hung stream, RSS growth < 1 GB p99 1.39 s, RSS +96 MB
gateway restarted under load, history rolled relay every worker's index within 1% of the engine's set ≤ 2.5 s, equal after the load the five busy workers exact 0.9 s after the registrations
gap beyond the relay's window relay snapshot resync, nothing unrecovered, no degraded rank out of routing 2.0 s into the cut, back 0.0 s
20 batches dropped on the wire relay refilled inside the relay, gateway sees no gap 21 recovered, 0 lost, gateway replay requests 0
publisher delayed / restarted relay lag seen, resync counted, routed to again lag 1.48 s, data_loss resync 0.2 s, routed again 1.2 s after the fault, nothing stale
router restart on the accelerator fleet relay, hardware serving ≤ 10 s, index rebuilt, p99 ≤ 2× steady first request routable 0.64 s, index at the engines' counts by 1.8 s
engine restart on the accelerator fleet relay, hardware detected, resynced, hit fraction kept SERVING again at kill+52 s, hit fraction 0.96 to 1.00 in every 10 s bucket

Every row passes, read the Measured column against the Pass rule, with the long rows continued below.

  • The mock rows ran on eight workers, routing state sampled at 100 ms, event timings at 200 ms, engine truth from the mock's admin API.
  • The relay rows used a 100-batch history so a resubscription from zero takes the snapshot path.
  • 4 of 5 relay rows passed before the start replay and the thin-worker slice, the publisher-restart row failing until the thin-worker warm-up rule.
  • Engine restarted, empty cache: re-admitted 0.8 s after the port answered, stream up at 1.0 s, 1,422 batches applied after.
  • Publisher restarted: the index fell 1,024 → 22 at the resync, 29 → 235 within the 20 s hits-only phase and 1,032 after the trickle.
  • All 45 hits-only requests reached the emptied worker through the diversion, hit rate 92.1% before and 82.7% after against oracles 94.5 and 84.8%.
  • Relay variant: 45 of 45 diverted, index 22 → 225 and regrown to 1,049, diverted hits 0.6% of post-fault requests, no hung streams.
  • On the two-hour replay c3 the diversion fired once: the warm-up rule ended at a 1,024-block growth cap the emptied worker passed within seconds.
  • Fixed in this series, a thin worker stays warming until the thinness ratio and the cap bounds the age rule alone.
  • On a 15-minute replay with a publisher restart, memberships recover 262,127 → 248,450 within the minute, against 230,895 and flat before the fix.
  • The emptied worker refills 58% of its pool in 30 s, 92% in 60 s at a 0.9 ratio.
  • Past the ratio it idles: the engine publishes a block only when it stores it, so the shared heads never reach the index.
  • The drill on the fixed gateway diverts 47 of 47 hits, index 8 → 253, the trickle reaching 0.46 of the fleet's level, 0.20 before.
  • Gateway restarted with the history rolled: the near-idle workers were exact 0.8 s after their relays primed from the engines' replay.
  • Gap beyond the window: OUT_OF_RANGE then a snapshot, the index equal to the engine's set, 1,468 = 1,468 under load.
  • Router restart: four Qwen3-8B vLLM workers at 16.6 req/s, relay window rolled, SIGKILL at +60 s, the only failures the 88 requests in flight.
  • Process up 0.02 s, health 0.36 s, workers re-registered 0.38 s, snapshot resyncs at 0.8 s, engine counts 42,251 / 42,252 / 42,254 / 29,663.
  • TTFT p50 / p99 142 / 734 ms steady, 153 / 746 over the first 30 s, cached_tokens credit 1.000 after and 0.986 steady.
  • The first-30-s p99 ratio is 1.02 on one window and 0.93 against the 60 s before the kill, with a budget of 2.0.
  • Engine restart on the fleet: 0.6B workers, one killed and relaunched on the same ports, the stream error and reconnect within 6 s.
  • The restart was detected at the first reconnect, with an out_of_range resync at kill+20 s and TTFT p50 17 to 19 ms throughout.

Simulator against hardware

The mock agrees with the accelerator fleet within 15% on 18 of 28 metrics at the full pool, and not yet on the restricted pool's tails.

Scenario Policy goodput req/s within SLO TTFT p50 ms TTFT p90 TTFT p99 prefix reuse
full pool (676,128 KV tokens per worker) cache_aware 14.32 / 15.31 (0.94) 86.8 / 92.5% (0.94) 175 / 156 (1.12) 506 / 440 (1.15) 1,075 / 864 (1.24) 0.420 / 0.388
full pool round_robin 13.81 / 14.29 (0.97) 83.0 / 85.7% (0.97) 199 / 197 (1.01) 408 / 444 (0.92) 702 / 973 (0.72) 0.39 / 0.39
restricted pool (12,000 blocks per worker) round_robin 13.80 / 11.88 (1.16) 82.9 / 71.5% (1.16) 200 / 245 (0.82) 407 / 1,838 710 / 4,787 0.370 / 0.363
restricted pool cache_aware 12.73 / 7.04 (1.81) 77.4 / 47.8% 212 / 821 1,342 / 13,217 5,970 / 24,479 0.379 / 0.366

Each cell is mock / hardware with the ratio in brackets, same Mooncake rows 0-1,999 at 3×, same replayer, four workers, means of 3 on both sides.

  • At the full pool 21 of 28 metrics are within 25% and the policy ordering on goodput and within-SLO is reproduced.
  • The systematic miss is the decode step, ITL 1.2 to 1.75× on the mock.
  • The fleet pays for the small pool where the mock does not, and its cache_aware collapse does not happen on the mock.
  • On the fleet vLLM announces removals late, 180k to 590k per worker per run, while the mock's index stays exact at hit/oracle 0.998.
  • Three mock causes were found and fixed on the way: head-first eviction, chunk-only admission, and the decode fit read against the configured pool.

Routing decision cost

At 128 workers a decision costs 8.9 µs at p50 and 15.9 µs at p99 with the chain index, inside the contract's 20 µs minimum and above its 10 µs goal.

Pool Backend decision p50 / p99 (µs), first measurement after bounding the selection stage to the deepest holders
128 workers chain index 29.6 / 55.5 (lookup 3.5 / 9.3, 12% of the decision) 8.9 / 15.9
128 workers positional 39.7 / 111.4 (lookup 9.2 / 94.5, 35%) 16.6 / 87.5
1,024 workers chain index 283.9 / 499.2 (3,465 decisions/s achieved) 24.4 / 35.6 (9,993 decisions/s)
1,024 workers positional 333 / 591 62 / 170

Release build, one pinned core, 10,000 decisions per second sustained, requests p50 1,344 and p99 8,288 tokens.

  • 88% of the chain index's decision was the selection stage's scan over the healthy workers plus hashing.
  • With shared chat-template heads the index names nearly every worker as a holder, so bounding the holders and the ties is what mattered.
  • At 1,024 workers the index lookup is 47% of the decision and is the next lever.
  • The same-source pair after the index fixes shows no regression in decision or lookup cost, its p99 in the Goals table.
  • A later tuning reads a run's holder table only when the coverage word leaves a question, and coverage words up to the highest worker id.
  • At 128 workers that reads lookup p50 / p99 3.62 / 9.47 µs against 3.87 / 10.43 before it, decision p50 8.4 to 8.8 µs.
  • The tuning restored c1's run structure but not its live lookup cost, p50 3.8 against 3.1 µs, the four-head A/B that attributes it running tonight.
  • 1,000 misses over 128 idle workers reach all, none above three times the 7.8 mean, and a cheaper worker still wins 1,000 of 1,000.

Policies

No policy met the end-to-end target against the external baseline, the default decision is unchanged, and protection on by default costs nothing measurable.

Load report, poll round_robin cache_aware least_load reference cost (external baseline) cache-aware-balanced (measured, not in the tree)
mock's full report, 10 s 144 / 2,179 / 305, 12.14, 0.720, 0.792, 1.03 180 / 4,436 / 423, 11.64, 0.702, 0.962, 3.08 219 / 2,577 / 364, 9.57, 0.572, 0.799, 1.36 163 / 1,960 / 280, 12.99, 0.765, 0.881, 1.13 234 / 3,633 / 467, 8.76, 0.525, 0.819, 1.22
vLLM-like report, 10 s 152 / 2,225 / 313, 12.13, 0.719, 0.786, 1.03 177 / 3,242 / 334, 13.46, 0.795, 0.986, 3.08 174 / 1,794 / 273, 12.51, 0.736, 0.802, 1.06 171 / 1,933 / 290, 12.76, 0.753, 0.882, 1.16 182 / 1,980 / 284, 12.44, 0.734, 0.828, 1.09
vLLM-like report, 1 s 145 / 2,068 / 301, 12.30, 0.728, 0.781, 1.02 160 / 2,027 / 269, 13.87, 0.824, 0.994, 3.09 145 / 2,097 / 289, 12.49, 0.739, 0.784, 1.04 169 / 1,964 / 286, 12.92, 0.760, 0.860, 1.23 134 / 1,915 / 259, 13.36, 0.787, 0.953, 1.06
hot prefix 1, 10 s: half the requests share one 3,072-token head, reuse 0.494 → 0.529 132.3 / 2,079, 12.67, 0.748, 0.875, 1.04 149.6 / 1,963, 13.49, 0.802, 0.993, 2.19, hottest worker 27% of requests 157.7 / 1,842, 12.84, 0.755, 0.876, 1.07 154.2 / 1,926, 13.28, 0.781, 0.923, 1.20 v2: 147.1 / 1,865, 13.43, 0.789, 0.972, 1.10
hot prefix 2, 10 s: 30% share one of four heads, reuse 0.494 → 0.513 146.7 / 2,135, 12.29, 0.727, 0.870, 1.03 163.3 / 2,061, 12.99, 0.768, 0.995, 2.22, hottest worker 29 to 31% 172.9 / 1,934, 12.52, 0.737, 0.877, 1.08 165.2 / 1,987, 12.99, 0.762, 0.928, 1.14 v2: 161.7 / 1,988, 12.98, 0.765, 0.957, 1.11
KV pressure, 12,000 blocks at 1.5×, loads up to 10 s stale 160.7 / 8,764 / 530, 6.15, 0.712, 0.893, 1.10, 4 195.3 / 9,622 / 630, 6.29, 0.736, 0.978, 2.70, 7 221.7 / 6,000 / 471, 5.33, 0.625, 0.900, 1.14, 4 186.9 / 3,326 / 354, 6.20, 0.720, 0.941, 1.19, 2 v2: 186.2 / 2,557 / 339, 6.27, 0.725, 0.955, 1.15, 0
KV pressure, fresh pushed load records 170.8 / 1,962 / 288, 6.81, 0.788, 0.990, 2.48, 0 176.9 / 2,218 / 322, 6.44, 0.746, 0.944, 1.14, 0 v2: 179.0 / 1,960 / 303, 6.36, 0.736, 0.962, 1.12, 0

8 mock workers at 3×, protection at token usage 0.8, means of 3, cells TTFT p50 / p99 / mean ms, goodput req/s, within SLO, hit/oracle, balance (hot-prefix rows: vLLM-like loads, no mean, balanced v2). The KV-pressure rows run at 1.5× with full load reports, oracle prefix reuse 0.340 at that budget, and end with preemptions per run. The reference-cost column is the external baseline, measured with bench-only binaries that are not in this PR.

Pool round_robin cache_aware least_load cache-aware-balanced (measured, not in the tree) cache_aware + token usage 0.8 cache_aware + waiting 8 balanced + tu 0.8 balanced + wq 8
full 14.48, 86.7%, 40.7%, 188-380-778, 0.392, 0.774 14.80, 88.6%, 45.2%, 174-359-771, 0.449, 0.887 14.26, 85.9%, 37.1% 14.38, 86.6%, 41.1% 14.73, 88.5% 14.74, 88.4% 14.39 14.41
12,000 blocks (7% of the working set), 128 sequences 14.19, 85.2%, 38.1%, 205-439-995, 0.368, 0.727 14.26, 86.1%, 38.9%, 184-379-932, 0.381, 0.753 14.11, 84.8%, 33.9% 13.93, 84.4%, 36.2% 14.11 (one worker at KV ≥ 0.95 for 10 s, p99 1,136) 14.00 (p99 1,310) 14.23 (p99 784) 14.21 (p99 860)

Accelerator fleet, four Qwen3-8B vLLM workers at 16 req/s, 1 s load poll, means of 3 at the full pool, single runs at 12,000 blocks, cells goodput, within SLO, strict SLO, TTFT p50-p90-p99 ms, prefix reuse, hit/oracle.

Policy protection on (default) protection off
round_robin 142 / 2,172, 12.24 144 / 3,553, 12.03
cache_aware 174 / 1,781, 14.11 172 / 1,822, 14.15
least_load 182 / 1,926, 12.24 181 / 2,030, 12.38
reference cost (external baseline) 163 / 1,896, 12.98 164 / 1,881, 12.91
cache-aware-balanced v2 (not in the tree) 159 / 1,845, 13.19 164 / 1,857, 12.97

Protection default against the opt-out on the mock with vLLM-like loads at the 10 s poll, 2 to 3 runs per cell, TTFT p50 / p99 ms and goodput req/s, every difference inside run-to-run noise.

  • Plain affinity did not collapse on the hot prefixes: the count-pressure gate replicates the hot chain onto three or four workers within a minute.
  • At four hot chains the gate concentrates them on two workers inside the SLO.
  • There the cost functions buy a flat fleet for the same goodput and 5% off the p99, v2 against the reference inside the noise.
  • A hot head opens no gap for the cost functions to close, so T5's axis on this fleet is not prefix heat.
  • The 12,000-block point at 3× does not discriminate, so the 1.5× KV-pressure point replaced it.
  • There every policy collapses alike: TTFT p50 29.6 to 31.6 s, goodput 0.92 to 1.61 req/s, 47 to 57 preemptions per run.
  • With fresh loads the default holds the tail under KV pressure and balanced buys only a flat fleet, which is why it leaves the tree.
  • The pressure finding was freshness: with records inside the poll interval the default still hoards, one worker taking 30 to 43% of requests, without preemptions.
  • The hottest worker's KV usage sits at 0.8 to 0.94, a 0.7 spread to the coolest, nothing preempted, so no KV-spread threshold is recommended.
  • Pushed load records hold the gateway's running-requests gauge within 0.06 of the engine's against 2.15 polled, staleness 0.81 against 10.55 s.
  • They cost 9 µs per event batch and 1 ms per idle heartbeat, gateway CPU +7% over the hour, no routing effect at that load.
  • Under normal load balanced costs about 3% goodput and 5 points of hit rate against the default, the full-pool rows above.
  • Balanced is SMG's own expected-wait score minus a capped relative prefix credit, and shares only the general overlap-credit idea with other routers.
  • On the hardware at the full pool the policies sit within noise of each other on goodput, because the fleet is not contended.
  • Earlier heads at the 10 s poll had cache_aware at 5 to 8.6 req/s at 12,000 blocks, hence the 1 s poll there.
  • least_load is the weakest row because it prices each worker's queue at that worker's own live generation rate.
  • Gateway CPU per request with protection on against off stays inside the same-binary control's spread over 54 cells.
  • HTTP at 64 tokens reads 0.616 against 0.589 ms, +4.6% with the control at +4.8%.
  • gRPC at 512 tokens reads 3.627 against 3.690 ms, −1.7% with the control at +1.2%.

Method

How the numbers were taken

  • One host with two sockets, a host-wide measurement lock held once per series, builds and the other lanes on other cores.
  • Both systems built and run the same day on the same cores with mimalloc, one fresh process per trial after 5 s of quiescence.
  • Two layouts: the comparison router's documented one, 64 event lanes and 128 query lanes, and a scaled one with more issuer cores for faster indexers.
  • Two memory settings, numactl --interleave=all as the published method and memory local to the socket, both reported, neither system changed for either.
  • A sustained point is bracketed with three fresh processes per rate, all three at 99% of offered with a valid generator, bisected to within 10%.
  • A series is 20 usable trials at that rate with interleaved same-binary controls, summarised as medians with bootstrap 95% intervals from 10,000 resamples.
  • No difference under 5% is called without the control pair showing a floor below it.
  • Cores are sampled around every trial, a foreign process above 5% of a core is recorded, above half a core the trial is replaced.
  • Across the scaled series 9 to 13 of 58 attempts per series were discarded, and the host carried other users' processes throughout.
  • One series was set aside for a wrong lane mask and rerun.
  • The hardware rows ran on four accelerators with warm engines, three labelled runs per row, the gateway and load generators on their own cores.

Caveats

  • The comparison harness has one query issuer, cannot drive two sockets, and its generator is invalid above about 1.5 to 2.1 billion block ops/s.
  • The two harnesses agree on sustained throughput to 0.2% at 107 M, and the SMG harness reads 7% lower at high rates.
  • The two-socket rows were taken on a loaded host and are labelled so.
  • The policy rows are unlocked on the mock, where the ordering is trustworthy and the magnitudes are not, and labelled per run on the hardware.
  • The adapter and patches that compile only inside the comparison tree stay outside this repository with the measurement scripts.

Goals, as rescoped on 2026-10-05, and where each track stands

Two tracks are fully met, two are met at the floor, and the end-to-end routing target is not, the minimum in brackets after each goal.

Track Goal (minimum) State Met
T1 indexer throughput ≥ 2× the comparison indexer's sustained block ops/s at equal cores, both memory settings (≥ 1.5×) 1.44× interleaved, 1.50× local memory, 1.37× on the scaled layout, capacity 1.53× floor met under local memory, missed by 4% under the published setting, goal not met
T2 indexer latency lookup p99 ≤ 1/3 of the comparison's at every sustained load, p50 ≤ 2 µs (p99 ≤ theirs) harness p99 3.6 to 4 µs against 9.9 to 14, p50 1.2 to 1.3 µs, the in-gateway churn growth fixed by children at any offset (churn table) met in the harness, flat on churn runs c1 and c3 and on s8
T3 index memory ≤ 1/10 of the comparison's bytes per indexed block at 128 workers, RSS flat over 24 h (≤ 1/5) 1/14.8 of the comparison's bytes per block, idle allocated flat across the soaks met, idle RSS flat over s8's six hours, the 24 h figure still owed
T4 routing accuracy hit rate ≥ 0.98 of the oracle, predicted = actual cached tokens, convergence ≤ 500 ms (≥ 0.95, ≤ 1 block, ≤ 1 s) hit/oracle 0.994 to 0.999 across backends and soaks, credit = cached_tokens 385 of 385 on the mock and 1.000 over 508 on hardware, convergence 0.6 s after a gap and 0.8 to 1.8 s after a restart, s7's twelve-hour mean 0.979 under the refill defect fixed here accuracy met, the 500 ms convergence not demonstrated
T5 end-to-end latency mean TTFT ≥ 40% lower and goodput ≥ 25% higher than the comparison cost function at 3× Mooncake, reuse ≥ 0.52 (25% / 15%) best candidate −21% TTFT p50 and +3.4% goodput at a 1 s poll, −3.2% / +1.5% at 10 s, −3.3% / +1.1% on the hot prefix, hardware within 3%, oracle reuse 0.39 to 0.51, under KV pressure balanced beats our own default only with stale loads (mean −46%, p99 −73%) not met against the external baseline on any axis
T6 routing decision cost ≤ 10 µs p99 per request at 10k req/s per core (≤ 20 µs) 55.5 µs p99 at 128 workers before the selection bound, 15.9 after it, 16.7 to 17.1 after the index fixes minimum met, goal not met
T7 scaling linear 8 to 128 workers and 1 to 64 lanes, no cliff across sockets (within 20% of linear) lanes 8/16/32/64 → 120/228/420/398 M before the pool, 1,062 M with it. One socket 1,234 M, series 1,229.8 M [1,224.6, 1,230.0] at 47 ns per block. Two sockets, two shards, 745.5 M, series 720.4 M [697.1, 734.8] at 55 / 56 ns per block, 2.3% duplicated, 0.6× one socket because one preempted lane fails the keep-up rule per-block efficiency met, throughput not
T8 fault tolerance every drill passes, index after recovery identical to a fresh snapshot (all pass) direct 12 of 12, relay 5 of 5, router restart on hardware with the index equal to the engines' sets met, by block counts and the engines' hit ratio
T9 hardware agreement simulation and hardware within 15% on T4 and T5 (within 25%) full pool 18 of 28 within 15% with the ordering reproduced, restricted pool within 16% for round robin only met at the full pool, not at the restricted pool

Decisions taken by the user during the loop

  • Goals rescoped on 2026-10-05: T1 from ≥ 10× to ≥ 2× with a 1.5× floor, T2 and T3 tightened, T7 unchanged.
  • The default routing policy does not change: cache_aware with affinity first, the count-pressure and KV-usage gates, and the overload protection.
  • cache-aware-balanced was measured, its rows stay in the tables, and it was removed from the tree, free to return in its own pull request.
  • The reference-cost policy leaves the tree, its rows stay in the tables as the external baseline from bench-only binaries no longer in the PR.
  • T5 is published as measured: not met against the external baseline on any axis, its KV-pressure win over our own default only with stale loads.
  • Nothing from the comparison router in the tree, not even tests.
  • The 10 s load poll is too stale for live load, so the servicer pushes its load on every batch, GetLoads only the fallback.
  • The positional indexer is deleted after the chain index's gateway soak and the user's default flip, not in this series.

Not in this series

  • The counting-allocator row after children at any offset, and the two-socket window with stealing lanes on a quiet second socket, with the per-socket lookup mode.
  • Soak s8's evidence is in, the default flip and the removal of the positional indexer and the run alias are the user's call.
  • A quiet-host series for the comparison router at its newer head, the restricted-pool hardware scrape, and the twelve-run simulator table on the final mock.
  • least_load's live-work pricing stays on a lane.
  • Run c5, the fixed refill on the two-hour replay at the default ratio, and s7's reading at 24 hours.
  • Per-rank identity end to end, the router-to-router index bootstrap and a gateway-side index dump are proposals.
  • The branch is rebased on main 529e3a3 with no conflicts left, 14 resolved over two merges of main into the loop branch.

Gates

The full gate ran on the loop tree this commit is cut from, and the quick gates on this head.

  • On the loop tree: nightly cargo fmt --all --check clean and pre-commit run --from-ref origin/main --to-ref HEAD with 15 hooks clean.
  • 3,011 tests passed, 0 failed, 5 ignored across 19 test binaries: kv-index, engine-servicer, mock-worker with main's capture tests, gateway library, routing, API and pushed-load tests.
  • Clippy in CI's configuration: the workspace at -D warnings, the leaf crates with every feature, the gateway's feature list, an x86_64 cross clippy of kv-index.
  • Python servicer tests with locally generated proto stubs: 539 passed, 18 skipped, 1 failed on a vllm.exceptions stub import that fails on main too.
  • On this head the leak gate reads 0 over every added line and the body, and the tree carries no binary files or engine captures.
  • The post-hoc set re-runs on this head after tonight's measurement window.

Provenance

Row family SMG head Comparison head Corpus Layout
This branch perf/kv-router-leap-pr at 4af6b8a, one commit from the loop tree 07d6187 on base 529e3a3, which is main, 14 conflicts resolved over two merges of main into the loop branch, the full gate on that loop tree
Scaled-layout series and brackets kv_index b4943d6 family in the comparison binary, SMG replayer 93876aa0 50bdb355f8, features mooncake and router-bench Mooncake trace, 128 workers, duplication 20, length factor 4, 128-token blocks, 2,446,195 operations, 320,105,993 block ops 8 event issuer cores, 4 query issuer cores, 52 lane cores
Quiet-host series and brackets, the cited ratios kv_index 9f9c7c0 in the same harness build 2b20fc1d35, indistinguishable from 50bdb355f8 at these points same corpus 5 issuer cores, 59 lane cores, interleaved and local memory
Load ladder and memory rows kv_index 9f9c7c0 and the pool-head builds before it 50bdb355f8 same corpus comparison layout, 3 interleaved trials per window
Decision-cost bench first row on the day's loop head, the bound 53577f3, the same-source pair af74524 generated corpus of 915,775 then 939,737 memberships one pinned core, 128 and 1,024 workers
Drills and policy tables binaries named by sha256 in the loop branch's protocol note, hardware rows 85f9d28 Mooncake rows 0-1,999 and 0-3,999 mock fleet of 8, accelerator fleet of 4

@coderabbitai

coderabbitai Bot commented Oct 6, 2026

Copy link
Copy Markdown

Important

Draft PR not reviewed

Draft PRs are not automatically reviewed by default.

  • Trigger a manual review

To automatically review draft PRs, update your CodeRabbit configuration:

reviews:
  auto_review:
    drafts: true
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Comment @coderabbitai help to get the list of available commands.

@github-actions github-actions Bot added documentation Improvements or additions to documentation python-bindings Python bindings changes dependencies Dependency updates grpc gRPC client and router changes benchmarks Benchmark changes tests Test changes protocols Protocols crate changes model-gateway Model gateway crate changes kv-index KV index crate changes labels Oct 6, 2026
@slin1237 slin1237 closed this Oct 6, 2026
@slin1237 slin1237 reopened this Oct 6, 2026
@slin1237
slin1237 force-pushed the perf/kv-router-leap-pr branch from 612e882 to 4dbf543 Compare October 6, 2026 19:23
@slin1237 slin1237 closed this Oct 6, 2026
@slin1237 slin1237 reopened this Oct 7, 2026
@slin1237
slin1237 force-pushed the perf/kv-router-leap-pr branch 2 times, most recently from da0bf1c to d148558 Compare October 7, 2026 03:43
@claude

claude Bot commented Oct 7, 2026 •

Copy link
Copy Markdown

👋 The PR description doesn't fully follow
PULL_REQUEST_TEMPLATE.md:

  • Missing header: ## Description
  • Missing header: ### Problem
  • Missing header: ### Solution
  • Missing header: ## Changes
  • Missing header: ## Test Plan

Please update the PR description so reviewers have the context they need.

… snapshots, and liveness-aware cache routing

The gateway's cache-aware routing keeps an index of which worker holds
which prefix of which prompt, fed by the engines' KV-cache event
streams. This series replaces the design of that index, the relay that
feeds it and the routing decisions around it.

The chain index stores the engines' block-hash chains run-length
compressed: a path-compressed trie of runs in an arena, one content hash
per position, one coverage bit per worker per run plus a table of
partial holders, children at any offset so a divergence inside a run
does not split it, lock-free readers that write nothing, writers that
lock one run at a time and descend without locks, recycled runs, arrays
and tables, a per-lane open-addressing block map, a lane pool that
schedules whole workers, and a sharded form (one index per NUMA node,
lookups unioned). It is exact by construction against a single-threaded
reference indexer added as a standing test, on recorded engine streams
and on seeded corpora with evicted middle blocks, inside the crate and
inside the gateway's own apply path. The positional indexer stays the
default behind `--kv-index {positional,chain}` until the chain index has
passed its gateway soak; its removal is a later change.

The servicers' KV-event relay decodes both wire layouts of both engines
into one model and normalizes per stream (tiers, cache groups, locality,
ownership, namespaces, bigram pages, one counter per drop reason),
forwards stores and removals one for one while the gateway counts
physical copies per tier and rank, reproduces both engines' chain hashes
so a misconfigured worker shows as a mismatch rate, subscribes every
data-parallel rank, keeps a bounded history so a resume is served from
it and a dropped batch is refilled from the engine's replay socket,
keeps the engine's live-block record and serves it as a state snapshot
to a subscriber whose cursor predates the history, starts at the
servicer's boot and primes itself from the engine's replay, reads a
publisher restart by rules that hold on every wire, and attaches the
engine's load to every batch with heartbeats while it is quiet, the load
poll kept only as a fallback. The vLLM servicers report queued uncached
token-work, generation throughput and hit rate from their own
bookkeeping.

The gateway admits events through one cursor per (worker, rank) with
bounded gap handling and snapshot resyncs; a liveness tracker beside the
health check vetoes a worker the transport knows is dead, partitioned or
wedged within about two seconds and re-admits it on first contact with a
closed circuit breaker; a thin or returned worker receives a slice of
cache-miss traffic; every request's end reaches the policy that placed
it through the worker's load guard; worker selection runs through a
cost-function selection layer whose default reproduces the existing
decision; worker overload protection is on by default as steering, never
as shedding. The mock worker becomes a vLLM-style engine with the
engines' ZMQ wires, fault hooks and engine truth, and its replay binary
scores every routing decision against the fleet's arrival-time oracle
and the engines' own cached-token counts. The stdout log sink no longer
blocks runtime threads, and jemalloc purges on schedule.

Measured in the strongest open-source KV router's own benchmark binary,
both indexers behind the same lanes on the same host and the same cores,
20 fresh-process trials per point with interleaved same-binary controls
and bootstrap intervals: the chain index sustains 1.44x that router's
block operations per second under the published memory policy and 1.50x
with node-local memory, at about a third of its lookup p99 and about one
fifteenth of its bytes per indexed block; in the gateway it matches the
positional path's hit rate with the lookup's p99 fifty times lower, and
every fault drill of the recovery protocol passes in both topologies.
Design: crates/kv_index/README.md. Benchmarks:
crates/kv_index/benches/README.md.

Signed-off-by: Simo Lin <25425177+slin1237@users.noreply.github.com>
@slin1237
slin1237 force-pushed the perf/kv-router-leap-pr branch from d148558 to 4af6b8a Compare October 7, 2026 05:36
pub fn from_gateway_config(config: &RouterConfig) -> Self {
if !config.worker_overload_protection {

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔴 Important: This early return breaks the Python launcher, which this PR didn't update. bindings/python/src/smg/router_args.py still defaults to worker_overload_protection: bool = False and worker_overload_waiting_requests / worker_overload_token_usage = None, and Router.from_args passes those values to _Router(**args_dict). The new pyo3 signature defaults (Some(8), Some(0.8), true) therefore never apply, which causes two problems:

  1. python -m smg.launch_router (no overload flags) keeps protection off. The Rust CLI now turns it on, so the two entry points disagree.
  2. --worker-overload-waiting-requests 16 without --worker-overload-protection used to enable protection ("either threshold set on its own enables protection"). Now it reaches this branch with worker_overload_protection == false and the user's explicit threshold is silently dropped.

The router_args.py help text also still describes the old 0.9 / "unset disables" behavior, and worker_overload_shed has no Python flag.

Fix: update RouterArgs (defaults 8 / 0.8 / True, add --disable-worker-overload-protection and --worker-overload-shed, update the help text) and test_arg_parser.py. Alternatively, treat an explicit threshold as enabling protection, as before.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

benchmarks Benchmark changes dependencies Dependency updates documentation Improvements or additions to documentation grpc gRPC client and router changes kv-index KV index crate changes model-gateway Model gateway crate changes protocols Protocols crate changes python-bindings Python bindings changes tests Test changes

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant