Skip to content

feat(rust): Laguna decode + batched prefill on dev (re-landed from #69/#70)#73

Merged
TheTom merged 5 commits into
devfrom
feat/laguna-dev
Jul 24, 2026
Merged

feat(rust): Laguna decode + batched prefill on dev (re-landed from #69/#70)#73
TheTom merged 5 commits into
devfrom
feat/laguna-dev

Conversation

@TheTom

@TheTom TheTom commented Jul 24, 2026

Copy link
Copy Markdown
Contributor

What

Re-lands Laguna-S-2.1 decode (#69) and Laguna batched prefill + weight cache + dtype-cache fix (#70) onto dev — they were merged into tom/feat/cuda-hip-vulkan-backends by mistake. Cherry-picked, swept to the Butter/Iron names.

Stacked on #71 (base = rename/butter-branding). Merge #71 first; this then retargets to dev. Split out of #71 per Tom to keep the rename PR pure.

FYI @ekryski.

Notes for review

  • Both PRs carried every op they use in their own diffs — no extra ops ported, none of the other 200 backends-branch commits included.
  • CUDA marlin trait methods from feat(rust): Laguna batched prefill, 156 to 559 pp tok/s, plus weight cache and dtype-cache fix #70 (moe_marlin_gemm/marlin_repack/marlin_build_routing) fall back to the default "unsupported" stub — their CudaDevice runtime methods aren't in thewafflehaus/iron@dev yet (implementations recovered, being validated on the 5090; iron-side PR to follow, then these get un-stubbed). Laguna itself never calls them — its W4A16 grouped GEMM dispatches raw kernels directly; only the NemotronH path is affected and it errors cleanly.
  • Two pre-existing dead-code test references removed (bench_smallmodel_fuse_slicecast, add_rms_norm_f16norm) — they broke the CUDA test-crate build on the source branch too.

Verification (M5 Max + agents)

@TheTom
TheTom requested a review from ekryski July 24, 2026 01:43
@github-actions github-actions Bot added the feature New feature or capability label Jul 24, 2026
TheTom added 4 commits July 24, 2026 11:12
#69) (renamed)

* sync(rust): land in-flight CUDA engine work from the GB10 working tree

Carries the uncommitted engine state the GB10 box has been running:
CUDA-graph capture trait plumbing (begin/end capture, graph_launch),
fp8 projection microbench and MoE grouped-MMA test updates, and the
device trait additions they depend on. Precedes the Laguna port commits
that build on these interfaces.

* feat(laguna): Laguna-S-2.1 decode on the CUDA engine

Full single-stream decode for the 117.55B/8.14B-active hybrid MoE:
sigmoid top-10 router with score-correction bias and shared expert,
per-head softplus attention gate, per-head QK RMSNorm, period-4
full/sliding-window attention with a 512-slot ring KV cache, YaRN
partial-rotary rope on full layers (scaling read from GGUF metadata,
this export is a 256K/factor-32 checkpoint), Q8 LM head, and a
Q4-requantized weight load straight from GGUF.

New ops: rope_yarn_partial (+posbuf variant), gate_softplus_mul_perhead,
argmax_f32_device, sdpa_decode_nbuf (n_kv from a device buffer),
write_u32, moe_gather_q4 and moe_gather_q4_swiglu wrappers (kept as
always-error stubs in this rebrand — the iron_moe_gather_q4 /
iron_moe_gather_q4_swiglu kernels aren't in thewafflehaus/iron@dev yet).

Decode runs eagerly or as a captured CUDA graph (BUTTER_LAGUNA_GRAPH=1,
warmup-capture-replay via LagunaDecodeCtx with a persistent workspace),
with optional fused MoE gate+up (BUTTER_LAGUNA_FUSE=1) and micro-fusions
(BUTTER_LAGUNA_MICRO=1: concatenated QKVG projection, shared expert via
the fused gather, o_proj residual accumulate).

GB10 single-stream greedy: 21.2 tok/s reference baseline, 36.5 tok/s
here (graph+fuse+micro), greedy tokens verified identical to the reference
implementation. Env-gated integration smoke: BUTTER_LAGUNA_GGUF plus
laguna tests in wh-butter-cuda (unit kernels run without weights).
…cache and dtype-cache fix (#70) (renamed)

* fix(cuda): key compiled-module and shared-size caches by kernel dtype signature

The ops layer caches kernel IR by (name, dtype) but the backend cached
compiled modules by bare kernel name, so the first dtype to touch a name
won and every other dtype silently ran the wrong binary: wrong element
stride, out-of-bounds reads, and in the shrinking-stride direction silent
corruption. Found when the prefill path's first f16 gather inherited the
decode path's f32 module. Shared-memory sizing had the same hazard.

* feat(laguna): batched prefill with tensor-core projections, fused MoE, CTA scheduling

Chunked multi-token prefill (default chunk 2048): batched YaRN/plain rope,
per-query windowed varlen attention with a window-aware KV-block skip,
linear sliding-window scratch compacted into the decode ring, grouped-GEMM
MoE with on-device descriptors, Marlin W4A16 tensor-core dense projections
(concatenated QKV, o_proj, dense FFN, shared expert), fused gate+up expert
stacks, and grouped-GEMM CTA scheduling (descending-size expert order plus
N-banded CTA order for weight L2 reuse, default on).

Correctness gates: prefill-then-decode greedy continuation byte-identical
to decode-only; last-token argmax matches the reference oracle; kernel
unit tests for the batched rope, batched gate, windowed varlen skip, fused
swiglu gather, and scheduling A/B on skewed synthetic groups.

GB10 single-stream: prefill 559 tok/s at 2048 (was 156 at first light),
504 at 8192; decode unchanged at 36.3 via graph replay. The reference C++ engine
on identical weights and box: 663 and 660 stock.

* feat(laguna): on-disk weight cache and windowed parallel conversion

Content-keyed cache of every converted engine-format weight blob (per
tensor artifact, keyed by format version, source GGUF identity, and the
load-shaping env flags), written on first conversion and mmap-read on
later loads. Conversion itself runs rayon-parallel over a bounded window
of layers (full parallelism held tens of GB of transients and got
OOM-killed on the shared 128GB). Warm reload: 56s, down from ~7.5
minutes; cache hits and misses are reported at load end.

Also carries the comment hygiene sweep across the Laguna files (dash
style, neutral phrasing for external references) and the stale
decode-only module doc fix.
…gates

Both cherry-picked commits (#69, #70) brought every op they need with
them — no separate ops-porting commit was required. These are gate-driven
fixups only:

- wh-butter-cuda/src/imp.rs: drop the moe_marlin_gemm/marlin_repack/
  marlin_build_routing pass-throughs PR #70 added — they call through to
  methods on wh_iron_runtime::CudaDevice that don't exist in
  thewafflehaus/iron@dev (confirmed: no `marlin` symbol anywhere in that
  repo). Falls back to the wh-butter-core default "unsupported on this
  backend" stub (already present, unaffected) instead of failing to
  compile. Same root cause as the moe_gather_q4/moe_gather_q4_swiglu stubs
  already documented in wh-butter-ops.
- wh-butter-cuda/tests/all_models.rs: dropped a `smallmodel_fuse_slicecast_ab`
  test calling `wh_butter_modeltests::bench_smallmodel_fuse_slicecast` —
  confirmed that function doesn't exist anywhere in wh-butter-modeltests
  on the original tom/feat/cuda-hip-vulkan-backends branch either
  (pre-existing dead reference predating this cherry-pick, not something
  it introduced).
- wh-butter-cuda/tests/f16norm_f32in.rs: dropped — same pre-existing-bug
  class, references a `add_rms_norm_f16norm` op that was never
  implemented on the source branch.
- rust/Cargo.lock: regenerated via `cargo build` (not hand-edited).
BUTTER_MARLIN_FIXTURE_DIR selects the reference-fixture dir; unset gives a
quiet [skipped] instead of unwrapping a hardcoded /tmp path (repo convention:
fixture/model paths come from env vars, absent -> skip).
@TheTom
TheTom changed the base branch from rename/butter-branding to dev July 24, 2026 16:13
@TheTom
TheTom force-pushed the feat/laguna-dev branch from 6498862 to b191b62 Compare July 24, 2026 16:13
@TheTom
TheTom marked this pull request as ready for review July 24, 2026 16:22
iron@dev's iron_strided_col_copy takes 5 bindings and has no internal idx
guard (the guarded 6-binding variant only exists on an un-landed kernels
feature branch). Drop the extra total binding and cover exactly s*width
threads, using the largest power-of-two block (<=64) dividing the total so
Laguna's non-64-aligned gate-column shapes stay correct.
@TheTom
TheTom merged commit c57e320 into dev Jul 24, 2026
6 checks passed
@TheTom
TheTom deleted the feat/laguna-dev branch July 24, 2026 17:29
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

feature New feature or capability

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant