Repository navigation
Conversation
|
Important Draft PR not reviewedDraft PRs are not automatically reviewed by default.
To automatically review draft PRs, update your CodeRabbit configuration: reviews:
auto_review:
drafts: true
Comment |
612e882 to
4dbf543
Compare
da0bf1c to
d148558
Compare
|
👋 The PR description doesn't fully follow
Please update the PR description so reviewers have the context they need. |
… snapshots, and liveness-aware cache routing
The gateway's cache-aware routing keeps an index of which worker holds
which prefix of which prompt, fed by the engines' KV-cache event
streams. This series replaces the design of that index, the relay that
feeds it and the routing decisions around it.
The chain index stores the engines' block-hash chains run-length
compressed: a path-compressed trie of runs in an arena, one content hash
per position, one coverage bit per worker per run plus a table of
partial holders, children at any offset so a divergence inside a run
does not split it, lock-free readers that write nothing, writers that
lock one run at a time and descend without locks, recycled runs, arrays
and tables, a per-lane open-addressing block map, a lane pool that
schedules whole workers, and a sharded form (one index per NUMA node,
lookups unioned). It is exact by construction against a single-threaded
reference indexer added as a standing test, on recorded engine streams
and on seeded corpora with evicted middle blocks, inside the crate and
inside the gateway's own apply path. The positional indexer stays the
default behind `--kv-index {positional,chain}` until the chain index has
passed its gateway soak; its removal is a later change.
The servicers' KV-event relay decodes both wire layouts of both engines
into one model and normalizes per stream (tiers, cache groups, locality,
ownership, namespaces, bigram pages, one counter per drop reason),
forwards stores and removals one for one while the gateway counts
physical copies per tier and rank, reproduces both engines' chain hashes
so a misconfigured worker shows as a mismatch rate, subscribes every
data-parallel rank, keeps a bounded history so a resume is served from
it and a dropped batch is refilled from the engine's replay socket,
keeps the engine's live-block record and serves it as a state snapshot
to a subscriber whose cursor predates the history, starts at the
servicer's boot and primes itself from the engine's replay, reads a
publisher restart by rules that hold on every wire, and attaches the
engine's load to every batch with heartbeats while it is quiet, the load
poll kept only as a fallback. The vLLM servicers report queued uncached
token-work, generation throughput and hit rate from their own
bookkeeping.
The gateway admits events through one cursor per (worker, rank) with
bounded gap handling and snapshot resyncs; a liveness tracker beside the
health check vetoes a worker the transport knows is dead, partitioned or
wedged within about two seconds and re-admits it on first contact with a
closed circuit breaker; a thin or returned worker receives a slice of
cache-miss traffic; every request's end reaches the policy that placed
it through the worker's load guard; worker selection runs through a
cost-function selection layer whose default reproduces the existing
decision; worker overload protection is on by default as steering, never
as shedding. The mock worker becomes a vLLM-style engine with the
engines' ZMQ wires, fault hooks and engine truth, and its replay binary
scores every routing decision against the fleet's arrival-time oracle
and the engines' own cached-token counts. The stdout log sink no longer
blocks runtime threads, and jemalloc purges on schedule.
Measured in the strongest open-source KV router's own benchmark binary,
both indexers behind the same lanes on the same host and the same cores,
20 fresh-process trials per point with interleaved same-binary controls
and bootstrap intervals: the chain index sustains 1.44x that router's
block operations per second under the published memory policy and 1.50x
with node-local memory, at about a third of its lookup p99 and about one
fifteenth of its bytes per indexed block; in the gateway it matches the
positional path's hit rate with the lookup's p99 fifty times lower, and
every fault drill of the recovery protocol passes in both topologies.
Design: crates/kv_index/README.md. Benchmarks:
crates/kv_index/benches/README.md.
Signed-off-by: Simo Lin <25425177+slin1237@users.noreply.github.com>
d148558 to
4af6b8a
Compare
| pub fn from_gateway_config(config: &RouterConfig) -> Self { | ||
| if !config.worker_overload_protection { |
There was a problem hiding this comment.
🔴 Important: This early return breaks the Python launcher, which this PR didn't update. bindings/python/src/smg/router_args.py still defaults to worker_overload_protection: bool = False and worker_overload_waiting_requests / worker_overload_token_usage = None, and Router.from_args passes those values to _Router(**args_dict). The new pyo3 signature defaults (Some(8), Some(0.8), true) therefore never apply, which causes two problems:
python -m smg.launch_router(no overload flags) keeps protection off. The Rust CLI now turns it on, so the two entry points disagree.--worker-overload-waiting-requests 16without--worker-overload-protectionused to enable protection ("either threshold set on its own enables protection"). Now it reaches this branch withworker_overload_protection == falseand the user's explicit threshold is silently dropped.
The router_args.py help text also still describes the old 0.9 / "unset disables" behavior, and worker_overload_shed has no Python flag.
Fix: update RouterArgs (defaults 8 / 0.8 / True, add --disable-worker-overload-protection and --worker-overload-shed, update the help text) and test_arg_parser.py. Alternatively, treat an explicit threshold as enabling protection, as before.
One squashed commit: the loop branch's tree without its lab notes, which stay on the loop branch with the full history and the scoreboard. Heads, base and corpus are in the Provenance table at the end.
What
The gateway's KV-aware routing gets a new index, a lossless event relay and liveness-aware routing, all exact against a reference.
--kv-index {positional,chain}, the positional indexer stays the default until its soak, andrunis a deprecated alias ofchain.--worker-stall-secs 2and--worker-wedge-secs 3.Why
The index was the bottleneck, the relay lost events, and failures were noticed late.
Benchmark publication
On the same cores of one host, the chain index sustains 1.44× the comparison indexer's block operations per second under the published memory setting and 1.50× with local memory, at a third of its lookup p99.
Same binary, equal cores, scaled layout
The chain index sustains 1.37× and reaches 1.53× the capacity of the comparison indexer on 52 lane cores, with p50 2.7× and p99 3.5× lower.
Published memory setting, series medians with 95% intervals, the sustained column is the headline and capacity a within-harness figure.
Both systems at their tops of stack on a quiet host, comparison layout
These are the rows the ratios cite, series median against series median: 1.44× interleaved and 1.50× with local memory, with 1.43× and 1.50× on the kept-up points.
--interleave=all(published method)--interleave=all--cpunodebind=0 --membind=0)Read the series column for the ratio and the lookup column for the latency gap, 59 lane cores for every system.
Lookup latency across the load ladder
At every window both systems sustain, the chain index answers in about 1.3 µs at p50 and about 4 µs at p99 against 3.3 to 3.5 and 12.6 to 14.1 µs.
Same binary, comparison layout, measurement cores, 3 interleaved trials per window, medians.
Memory per block
The chain index holds 1/14.8 of the comparison indexer's bytes per resident block in the benchmark's shape.
Counting allocator in the shared harness binary at the 3 s window, 2,096,883 resident blocks for every backend.
Exactness
All three indexers agree with the reference on every score, count and membership, and every defect found on the way was fixed and pinned by a test.
The two backends in the gateway
In the gateway the chain index matches the positional path's hit rate with the lookup p99 fifty times lower and less memory.
--kv-index chain)Mock fleet of 8 realistic workers, Mooncake rows 0-3,999 at 3×,
cache_aware, 3 interleaved runs each, medians, 0 errors, goodput and reuse inside the unlocked noise.s6 showed the lookup cost growing with churn at constant memberships, and children at any offset fixed it in this series.
The churn bench stores decode blocks under the prompt's last block, as engines do.
Fault drills
Every drill of the recovery protocol passes in both topologies, and a router restart on real engines is a 0.6 s outage with the index rebuilt in under two seconds.
wedged3.1 s into the pause, routable 0.0 s after resumedata_lossresync 0.2 s, routed again 1.2 s after the fault, nothing staleEvery row passes, read the Measured column against the Pass rule, with the long rows continued below.
OUT_OF_RANGEthen a snapshot, the index equal to the engine's set, 1,468 = 1,468 under load.cached_tokenscredit 1.000 after and 0.986 steady.out_of_rangeresync at kill+20 s and TTFT p50 17 to 19 ms throughout.Simulator against hardware
The mock agrees with the accelerator fleet within 15% on 18 of 28 metrics at the full pool, and not yet on the restricted pool's tails.
Each cell is mock / hardware with the ratio in brackets, same Mooncake rows 0-1,999 at 3×, same replayer, four workers, means of 3 on both sides.
cache_awarecollapse does not happen on the mock.Routing decision cost
At 128 workers a decision costs 8.9 µs at p50 and 15.9 µs at p99 with the chain index, inside the contract's 20 µs minimum and above its 10 µs goal.
Release build, one pinned core, 10,000 decisions per second sustained, requests p50 1,344 and p99 8,288 tokens.
Policies
No policy met the end-to-end target against the external baseline, the default decision is unchanged, and protection on by default costs nothing measurable.
8 mock workers at 3×, protection at token usage 0.8, means of 3, cells TTFT p50 / p99 / mean ms, goodput req/s, within SLO, hit/oracle, balance (hot-prefix rows: vLLM-like loads, no mean, balanced v2). The KV-pressure rows run at 1.5× with full load reports, oracle prefix reuse 0.340 at that budget, and end with preemptions per run. The reference-cost column is the external baseline, measured with bench-only binaries that are not in this PR.
Accelerator fleet, four Qwen3-8B vLLM workers at 16 req/s, 1 s load poll, means of 3 at the full pool, single runs at 12,000 blocks, cells goodput, within SLO, strict SLO, TTFT p50-p90-p99 ms, prefix reuse, hit/oracle.
Protection default against the opt-out on the mock with vLLM-like loads at the 10 s poll, 2 to 3 runs per cell, TTFT p50 / p99 ms and goodput req/s, every difference inside run-to-run noise.
cache_awareat 5 to 8.6 req/s at 12,000 blocks, hence the 1 s poll there.least_loadis the weakest row because it prices each worker's queue at that worker's own live generation rate.Method
How the numbers were taken
numactl --interleave=allas the published method and memory local to the socket, both reported, neither system changed for either.Caveats
Goals, as rescoped on 2026-10-05, and where each track stands
Two tracks are fully met, two are met at the floor, and the end-to-end routing target is not, the minimum in brackets after each goal.
cached_tokens385 of 385 on the mock and 1.000 over 508 on hardware, convergence 0.6 s after a gap and 0.8 to 1.8 s after a restart, s7's twelve-hour mean 0.979 under the refill defect fixed hereDecisions taken by the user during the loop
cache_awarewith affinity first, the count-pressure and KV-usage gates, and the overload protection.cache-aware-balancedwas measured, its rows stay in the tables, and it was removed from the tree, free to return in its own pull request.GetLoadsonly the fallback.Not in this series
runalias are the user's call.least_load's live-work pricing stays on a lane.Gates
The full gate ran on the loop tree this commit is cut from, and the quick gates on this head.
cargo fmt --all --checkclean andpre-commit run --from-ref origin/main --to-ref HEADwith 15 hooks clean.-D warnings, the leaf crates with every feature, the gateway's feature list, an x86_64 cross clippy of kv-index.vllm.exceptionsstub import that fails onmaintoo.Provenance
perf/kv-router-leap-prat 4af6b8a, one commit from the loop tree 07d6187 on base 529e3a3, which is main, 14 conflicts resolved over two merges of main into the loop branch, the full gate on that loop tree