Skip to content

[ExecuTorch][WebGPU] Replace broad QKV fusion with BK64 kernel - #21650

Merged
JCNTH merged 12 commits into
gh/JCNTH/106/origfrom
gh/JCNTH/107/orig
Aug 7, 2026
Merged

[ExecuTorch][WebGPU] Replace broad QKV fusion with BK64 kernel#21650
JCNTH merged 12 commits into
gh/JCNTH/106/origfrom
gh/JCNTH/107/orig

Conversation

@pytorchbot

Copy link
Copy Markdown
Collaborator

This PR was created by the merge bot to help merge the original PR into the main branch.
ghstack PR number: #21131 by @JCNTH
^ Please use this as the source of truth for the PR details, comments, and reviews
ghstack PR base: https://github.com/pytorch/executorch/tree/gh/JCNTH/107/base
ghstack PR head: https://github.com/pytorch/executorch/tree/gh/JCNTH/107/head
Merge bot PR base: https://github.com/pytorch/executorch/tree/gh/JCNTH/106/orig
Merge bot PR head: https://github.com/pytorch/executorch/tree/gh/JCNTH/107/orig

@diff-train-skip-merge

Pull Request resolved: #21131

The previously landed broad QKV fusion applied too widely and did not match the
BK64 schedule now used for the ordinary projections. This corrective diff
replaces it with a capability- and geometry-qualified BK64 kernel that fuses the
exact Llama Q/K/V projection triple at live M128, M508, and M512, packing the
constant weights and scales once and scattering the result into three distinct
planner-safe outputs. Outside the accepted shapes it switches atomically back
to the ordinary Steel and bicol routes, and it deletes the obsolete broad
shader and header so only one QKV path remains. Mirrors Vulkan
xplat/executorch/backends/vulkan/runtime/graph/ops/glsl/q4gsw_linear_gemm__w_4x8.glsl
for the per-projection GEMM; the three-output fusion itself is WebGPU-specific.

Key changes:
- Adds runtime/ops/quantized_linear/q4gsw_qkv_bk64.wgsl (+ generated header);
  removes the obsolete q4gsw_linear_gemm_qkv_fused shader and header.
- WebGPUGraph.{cpp,h}: geometry and capability predicate, one-time constant
  packing, distinct Q/K/V outputs, and atomic fallback to Steel/bicol.
ghstack-source-id: 411961443
@exported-using-ghexport

Differential Revision: [D113171749](https://our.internmc.facebook.com/intern/diff/D113171749/)
@pytorch-bot

pytorch-bot Bot commented Aug 7, 2026

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/21650

Note: Links to docs will display an error until the docs builds have been completed.

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@meta-cla meta-cla Bot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Aug 7, 2026
JCNTH and others added 11 commits August 7, 2026 10:50
Pull Request resolved: #21132

Materializing the full attention-weight matrix for Llama prefill is memory- and
bandwidth-heavy and does not scale to longer sequences. This adds a single-pass
online-softmax K16 causal-attention kernel for the exact Llama Hq32/Hkv8/G4/D64
geometry with fp16 KV, guarded by explicit adapter limits, so accepted prefill
shapes through S512 never materialize attention weights. S1 keeps the existing
FlashDecoding path, and unsupported geometry, storage, or capabilities fall back
to the materialized implementation. No Vulkan analogue (WebGPU-specific): the
Vulkan backend has only a materialized attention (compute the weights, then a
separate multi-pass softmax), with no online-softmax kernel.

Key changes:
- runtime/ops/sdpa/streaming_attention_k16_causal_bound.wgsl (+ generated
  header): the online-softmax single-pass causal kernel.
- Sdpa.cpp: exact Llama geometry and capability guards, S512 transition
  handling, and fallback to the materialized route.
ghstack-source-id: 411961446
@exported-using-ghexport

Differential Revision: [D113171738](https://our.internmc.facebook.com/intern/diff/D113171738/)
Pull Request resolved: #21133

Building the WebGPU backend for the browser previously required a manual
source-tree patch to pick the right WebGPU implementation. This makes the CMake
configuration select the emdawnwebgpu port for Emscripten builds while native
builds keep Dawn and their platform libraries, so the production kernel stack
builds for the browser unpatched. Build-system integration only; no model or
kernel behavior changes.

Key changes:
- CMakeLists.txt: choose emdawnwebgpu under Emscripten, Dawn otherwise.
- test/test_cmake_configuration.py: contract test asserting the selection.
ghstack-source-id: 411961450
@exported-using-ghexport

Differential Revision: [D113171750](https://our.internmc.facebook.com/intern/diff/D113171750/)
Pull Request resolved: #21134

With dynamic sequence positions, the backend could record only one SDPA route,
so a decode following a dynamic-position prefill could dispatch the wrong
kernel, and there was no way to confirm which kernel actually ran. This treats a
dynamic SymInt position (not just a dynamic sequence dimension) as sufficient to
record both SDPA routes, exposes active-kernel route attestation for correctness
and performance harnesses, and reads the timestamp-query gate per execution so
diagnostics can be enabled after module initialization. The attestation state
and query are compiled only under the WGPU_BACKEND_ENABLE_PROFILING build flag,
so production builds carry no additional state or cost, mirroring how the Vulkan
backend gates its QueryPool behind ET_EVENT_TRACER_ENABLED. No Vulkan analogue
for the routing fix (WebGPU runtime routing and observability).

Key changes:
- WebGPUGraph.{cpp,h}: expose the active-route attestation query (compile-gated
  behind WGPU_BACKEND_ENABLE_PROFILING) and read the timestamp-query gate per
  execution.
- WebGPUUtils.h, Sdpa.cpp: treat a dynamic SymInt position (not only a dynamic
  sequence dimension) as sufficient to record both SDPA routes.
ghstack-source-id: 411961454
@exported-using-ghexport

Differential Revision: [D113171745](https://our.internmc.facebook.com/intern/diff/D113171745/)
Pull Request resolved: #21135

**Add HuggingFace rotate-half RoPE and migrate RoPE to shared dispatch construction**

Qwen models pair the first and second halves of each head vector rather than adjacent elements. This adds the rotate-half operator with dynamic start-position updates, full-dimension Q/K handling, generated WGSL, and strict frequency-table bounds.

The shared RoPE handler also closes the dispatch-boilerplate review: it uses `graph.device()`, typed graph-owned parameter buffers, descriptor-driven bindings and pipelines, named validation and resize callbacks, generated shader-registry lookup, scoped ownership, and graph-owned workgroup recomputation. Interleaved shader math and output are unchanged.

Key changes:
- `rotary_embedding_hf.wgsl` and generated registry entry — one thread per pair for HuggingFace rotate-half.
- `RotaryEmbedding.cpp` — named validation, typed resize contexts, shared dispatch descriptors, and one initial/resize grid picker per route.
- Native tests — malformed input rejection, dynamic bounds, shrink/regrow reuse, and HF/interleaved lifecycle coverage.

Co-authored-with: Claude Code.
ghstack-source-id: 411961455
@exported-using-ghexport

Differential Revision: [D113171746](https://our.internmc.facebook.com/intern/diff/D113171746/)
Pull Request resolved: #21136

Qwen3's attention geometry differs from Llama's, and its KV cache is produced in
fp32 on the host but must be consumed in fp16 on the device. This adds guarded
Qwen3 K16 streaming online-softmax attention schedules — a Q16 schedule that is
the automatic default whenever the exact geometry and capability guards pass,
plus a Q32 candidate that is opt-in through the `sdpa_query_tile` runtime spec
(BackendOption) for future autotuning — together with the exact fp32-host to
fp16-device KV-cache boundary conversion. Selection requires the exact Qwen3
geometry, fp16 KV storage, adapter limits, a valid workgroup count, and an exact
2:1 byte ratio; the established Llama, materialized, and FlashDecoding routes
remain fallbacks. This builds on the HuggingFace rotate-half
RoPE operator. No Vulkan analogue (WebGPU-specific online-softmax attention;
Vulkan has only a materialized attention). It also makes the long generated WGSL
provenance and constant declarations format-stable and covers them with a
generator regression test.

Key changes:
- runtime/ops/sdpa/streaming_attention_qwen3_k16_causal_bound.wgsl and
  streaming_attention_qwen3_q32_k16_causal_bound.wgsl (+ generated headers): the
  Q16 and Q32 online-softmax Qwen3 kernels.
- Sdpa.cpp, WebGPUGraph.{cpp,h}: exact Qwen3 geometry and limit guards, Q16
  default route selection, and the fp32-host to fp16-device KV-cache conversion.
- WebGPUBackend.cpp: read the optional `sdpa_query_tile` runtime spec and thread
  it into graph build so the Q32 tile can be requested without a rebuild.
- scripts/gen_wgsl_headers.py (+ test_wgsl_codegen.py): format-stable generated
  headers with a regression test.
ghstack-source-id: 411961459
@exported-using-ghexport

Differential Revision: [D113171744](https://our.internmc.facebook.com/intern/diff/D113171744/)
Pull Request resolved: #21448

**Make WGSL publication rollback-safe without changing generated runtime output**

Generated headers and the registry were previously written incrementally, so later validation, interruption, or write failure could leave a partial generated tree. This change stages and validates the complete output set before sequential publication, rejects path/name/C++ symbol collisions, and restores replaced outputs on failure or interruption.

The family model follows Vulkan's YAML/template generation. The WebGPU publisher adds staging, backup, rollback, and temporary-file cleanup; it does not claim filesystem-atomic multi-file replacement.

Key changes:
- `gen_wgsl_headers.py` — stage, validate, publish, restore on failure, and clean temporary outputs.
- `test_wgsl_codegen.py` — cover collisions, orphan repair, interruption, rollback, and idempotence.

All generated shader bytes and runtime registry entries remain unchanged.

Co-authored-with: Claude Code.
ghstack-source-id: 411961465
@exported-using-ghexport

Differential Revision: [D113979658](https://our.internmc.facebook.com/intern/diff/D113979658/)
Pull Request resolved: #21449

**Generate typed `to_copy` variants from one byte-identical template**

The two conversion directions duplicated the same WGSL structure and could drift independently. This moves fp32-to-int32 and int32-to-fp32 into one typed family while preserving both expanded payloads. Mirrors Vulkan `backends/vulkan/runtime/graph/ops/glsl/view_convert_buffer.{glsl,yaml}`.

Key changes:
- `to_copy_convert.wgsl` and YAML — declare the two typed variants.
- Generated headers and codegen locks — preserve symbols, workgroups, registry entries, and payload hashes.
- Structural/native tests — lock both directions and the round trip.

Host routing and dispatch remain unchanged; tests expand fixture coverage.

Co-authored-with: Claude Code.
ghstack-source-id: 411961470
@exported-using-ghexport

Differential Revision: [D113979712](https://our.internmc.facebook.com/intern/diff/D113979712/)
Pull Request resolved: #21450

**Generate extrema and unary shader families**

The extrema reductions and ten no-parameter unary kernels duplicated shader skeletons that could drift independently. This consolidates amax/amin behind one extrema template and abs/cos/exp/hardswish/neg/round/rsqrt/sin/sqrt/tanh behind one unary template while preserving the generated runtime payloads.

Key changes:
- Generate amax/amin from one extrema manifest.
- Generate ten unary payloads from one operator-expression manifest.
- Lock expanded bytes, registry entries, delegation, and boundary numerics.

The attempted Unary lifecycle migration is intentionally not part of the stack: its performance campaign did not produce an authoritative passing result, so the Unary builder, interface, and activation/sigmoid call sites are restored to their pre-migration bytes.

Co-authored-with: Claude Code.
ghstack-source-id: 411961475
@exported-using-ghexport

Differential Revision: [D113979760](https://our.internmc.facebook.com/intern/diff/D113979760/)
…riants

Pull Request resolved: #21451

**Generate byte-identical logical and arithmetic binary variants from shared WGSL families**

Logical AND/OR and four arithmetic binary kernels duplicated shader skeletons and broadcast logic. This consolidates logical AND/OR behind one packed-Boolean family and minimum/pow/floor_divide/mul into the existing binary family, with a permanent mixed-rank broadcast contract.

Key changes:
- Generate logical AND/OR from one operator-token manifest.
- Generate minimum, pow, floor_divide, and mul beside the existing div/sub variants.
- Lock same-shape and mixed-rank expressions, exact payloads/workgroups, PTE delegation, and broadcast boundary cases.

No runtime C++ dispatch, bindings, pipeline construction, workgroups, or expanded shader payloads change. Four standalone WGSL inputs are removed, and future compatible variants require manifest entries instead of copied kernels. This follows the Vulkan binary-family pattern.

Co-authored-with: Claude Code.
ghstack-source-id: 411961479
@exported-using-ghexport

Differential Revision: [D113979789](https://our.internmc.facebook.com/intern/diff/D113979789/)
…ynamic resize

Pull Request resolved: #21483

**Preserve immutable Linear and Q4 embedding shader controls across dynamic resize through typed parameter authorities.**

**Problem**
The Linear resize hook dropped has_bias, and the Q4 embedding resize path could drop is_linear_weight. Both defects produced correct build-time outputs but silently changed semantics after a live shape update.

**Solution**
- Before: build and resize populated control words independently, allowing immutable shader state to reset.
- After: each build/resize pair calls one typed helper while recomputing only live counts and dispatch.

**Implementation**
- Linear.cpp — make_linear_params preserves the bias flag across vec4 and tiled routes.
- EmbeddingQ4gsw.cpp — EmbeddingLayout and make_embedding_params preserve nibble layout across resize.
- Mirrors Vulkan runtime/graph/ops/impl/Linear.cpp and EmbeddingQ4gsw.cpp, which carry immutable shader controls through dynamic dispatch.

**Constraints**
WGSL bytes, shader selection, bindings, uniform sizes, queue-write counts, dispatch formulas, and pipeline topology are unchanged. Production handler code is smaller because duplicated field population and scalar captures are removed.

Co-authored-with: Claude Code.
ghstack-source-id: 411961489
@exported-using-ghexport

Differential Revision: [D113992326](https://our.internmc.facebook.com/intern/diff/D113992326/)
This PR was created by the merge bot to help merge the original PR into
the main branch.
ghstack PR number: #21597 by
@JCNTH
^ Please use this as the source of truth for the PR details, comments,
and reviews
ghstack PR base:
https://github.com/pytorch/executorch/tree/gh/JCNTH/204/base
ghstack PR head:
https://github.com/pytorch/executorch/tree/gh/JCNTH/204/head
Merge bot PR base:
https://github.com/pytorch/executorch/tree/gh/JCNTH/203/orig
Merge bot PR head:
https://github.com/pytorch/executorch/tree/gh/JCNTH/204/orig
Differential Revision:
[D114936148](https://our.internmc.facebook.com/intern/diff/D114936148/)
@diff-train-skip-merge

cc @SS-JIA @manuelcandales @digantdesai @cbilgin

---------

Co-authored-by: Julian Ng-Thow-Hing <juliannth@meta.com>
Co-authored-by: Julian Ng-Thow-Hing <107437036+JCNTH@users.noreply.github.com>
@JCNTH
JCNTH requested a review from SS-JIA as a code owner August 7, 2026 17:50
@JCNTH
JCNTH merged commit 0f65c9a into gh/JCNTH/106/orig Aug 7, 2026
31 of 32 checks passed
@JCNTH
JCNTH deleted the gh/JCNTH/107/orig branch August 7, 2026 17:50
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants