[ExecuTorch][WebGPU] Route exact prefill shapes to the BK64 GEMM - #21649
Merged
Conversation
Pull Request resolved: #21130 Llama prefill drives the ordinary quantized-linear projections at a small set of fixed batch-row counts, and the generic Steel schedule leaves throughput on the table for them. This adds a frozen BK64 Steel quantized-GEMM route for the exact accepted Llama ordinary-projection shapes at live M128, M508, and M512, selected only when the capability and dynamic-route guards all pass. M511 and other prefill sizes stay on the generic Steel schedule and M1 stays on bicol decode, so the route fails closed outside its accepted shapes. Mirrors Vulkan xplat/executorch/backends/vulkan/runtime/graph/ops/glsl/q4gsw_linear_gemm__w_4x8.glsl (the 4-bit grouped-symmetric-weight GEMM), specialized here for the BK64 tile. Key changes: - runtime/ops/quantized_linear/q4gsw_steel_bk64.wgsl (+ generated header): the BK64-tiled q4gsw prefill GEMM kernel. - QuantizedLinear.cpp, WebGPUUtils.h: exact-shape predicate, capability guard, and dynamic re-entry into and out of the BK64 route. ghstack-source-id: 411961445 @exported-using-ghexport Differential Revision: [D113171739](https://our.internmc.facebook.com/intern/diff/D113171739/)
🔗 Helpful Links🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/21649
Note: Links to docs will display an error until the docs builds have been completed. This comment was automatically generated by Dr. CI and updates every 15 minutes. |
Pull Request resolved: #21131 The previously landed broad QKV fusion applied too widely and did not match the BK64 schedule now used for the ordinary projections. This corrective diff replaces it with a capability- and geometry-qualified BK64 kernel that fuses the exact Llama Q/K/V projection triple at live M128, M508, and M512, packing the constant weights and scales once and scattering the result into three distinct planner-safe outputs. Outside the accepted shapes it switches atomically back to the ordinary Steel and bicol routes, and it deletes the obsolete broad shader and header so only one QKV path remains. Mirrors Vulkan xplat/executorch/backends/vulkan/runtime/graph/ops/glsl/q4gsw_linear_gemm__w_4x8.glsl for the per-projection GEMM; the three-output fusion itself is WebGPU-specific. Key changes: - Adds runtime/ops/quantized_linear/q4gsw_qkv_bk64.wgsl (+ generated header); removes the obsolete q4gsw_linear_gemm_qkv_fused shader and header. - WebGPUGraph.{cpp,h}: geometry and capability predicate, one-time constant packing, distinct Q/K/V outputs, and atomic fallback to Steel/bicol. ghstack-source-id: 411961443 @exported-using-ghexport Differential Revision: [D113171749](https://our.internmc.facebook.com/intern/diff/D113171749/)
Pull Request resolved: #21132 Materializing the full attention-weight matrix for Llama prefill is memory- and bandwidth-heavy and does not scale to longer sequences. This adds a single-pass online-softmax K16 causal-attention kernel for the exact Llama Hq32/Hkv8/G4/D64 geometry with fp16 KV, guarded by explicit adapter limits, so accepted prefill shapes through S512 never materialize attention weights. S1 keeps the existing FlashDecoding path, and unsupported geometry, storage, or capabilities fall back to the materialized implementation. No Vulkan analogue (WebGPU-specific): the Vulkan backend has only a materialized attention (compute the weights, then a separate multi-pass softmax), with no online-softmax kernel. Key changes: - runtime/ops/sdpa/streaming_attention_k16_causal_bound.wgsl (+ generated header): the online-softmax single-pass causal kernel. - Sdpa.cpp: exact Llama geometry and capability guards, S512 transition handling, and fallback to the materialized route. ghstack-source-id: 411961446 @exported-using-ghexport Differential Revision: [D113171738](https://our.internmc.facebook.com/intern/diff/D113171738/)
Pull Request resolved: #21133 Building the WebGPU backend for the browser previously required a manual source-tree patch to pick the right WebGPU implementation. This makes the CMake configuration select the emdawnwebgpu port for Emscripten builds while native builds keep Dawn and their platform libraries, so the production kernel stack builds for the browser unpatched. Build-system integration only; no model or kernel behavior changes. Key changes: - CMakeLists.txt: choose emdawnwebgpu under Emscripten, Dawn otherwise. - test/test_cmake_configuration.py: contract test asserting the selection. ghstack-source-id: 411961450 @exported-using-ghexport Differential Revision: [D113171750](https://our.internmc.facebook.com/intern/diff/D113171750/)
Pull Request resolved: #21134 With dynamic sequence positions, the backend could record only one SDPA route, so a decode following a dynamic-position prefill could dispatch the wrong kernel, and there was no way to confirm which kernel actually ran. This treats a dynamic SymInt position (not just a dynamic sequence dimension) as sufficient to record both SDPA routes, exposes active-kernel route attestation for correctness and performance harnesses, and reads the timestamp-query gate per execution so diagnostics can be enabled after module initialization. The attestation state and query are compiled only under the WGPU_BACKEND_ENABLE_PROFILING build flag, so production builds carry no additional state or cost, mirroring how the Vulkan backend gates its QueryPool behind ET_EVENT_TRACER_ENABLED. No Vulkan analogue for the routing fix (WebGPU runtime routing and observability). Key changes: - WebGPUGraph.{cpp,h}: expose the active-route attestation query (compile-gated behind WGPU_BACKEND_ENABLE_PROFILING) and read the timestamp-query gate per execution. - WebGPUUtils.h, Sdpa.cpp: treat a dynamic SymInt position (not only a dynamic sequence dimension) as sufficient to record both SDPA routes. ghstack-source-id: 411961454 @exported-using-ghexport Differential Revision: [D113171745](https://our.internmc.facebook.com/intern/diff/D113171745/)
Pull Request resolved: #21135 **Add HuggingFace rotate-half RoPE and migrate RoPE to shared dispatch construction** Qwen models pair the first and second halves of each head vector rather than adjacent elements. This adds the rotate-half operator with dynamic start-position updates, full-dimension Q/K handling, generated WGSL, and strict frequency-table bounds. The shared RoPE handler also closes the dispatch-boilerplate review: it uses `graph.device()`, typed graph-owned parameter buffers, descriptor-driven bindings and pipelines, named validation and resize callbacks, generated shader-registry lookup, scoped ownership, and graph-owned workgroup recomputation. Interleaved shader math and output are unchanged. Key changes: - `rotary_embedding_hf.wgsl` and generated registry entry — one thread per pair for HuggingFace rotate-half. - `RotaryEmbedding.cpp` — named validation, typed resize contexts, shared dispatch descriptors, and one initial/resize grid picker per route. - Native tests — malformed input rejection, dynamic bounds, shrink/regrow reuse, and HF/interleaved lifecycle coverage. Co-authored-with: Claude Code. ghstack-source-id: 411961455 @exported-using-ghexport Differential Revision: [D113171746](https://our.internmc.facebook.com/intern/diff/D113171746/)
Pull Request resolved: #21136 Qwen3's attention geometry differs from Llama's, and its KV cache is produced in fp32 on the host but must be consumed in fp16 on the device. This adds guarded Qwen3 K16 streaming online-softmax attention schedules — a Q16 schedule that is the automatic default whenever the exact geometry and capability guards pass, plus a Q32 candidate that is opt-in through the `sdpa_query_tile` runtime spec (BackendOption) for future autotuning — together with the exact fp32-host to fp16-device KV-cache boundary conversion. Selection requires the exact Qwen3 geometry, fp16 KV storage, adapter limits, a valid workgroup count, and an exact 2:1 byte ratio; the established Llama, materialized, and FlashDecoding routes remain fallbacks. This builds on the HuggingFace rotate-half RoPE operator. No Vulkan analogue (WebGPU-specific online-softmax attention; Vulkan has only a materialized attention). It also makes the long generated WGSL provenance and constant declarations format-stable and covers them with a generator regression test. Key changes: - runtime/ops/sdpa/streaming_attention_qwen3_k16_causal_bound.wgsl and streaming_attention_qwen3_q32_k16_causal_bound.wgsl (+ generated headers): the Q16 and Q32 online-softmax Qwen3 kernels. - Sdpa.cpp, WebGPUGraph.{cpp,h}: exact Qwen3 geometry and limit guards, Q16 default route selection, and the fp32-host to fp16-device KV-cache conversion. - WebGPUBackend.cpp: read the optional `sdpa_query_tile` runtime spec and thread it into graph build so the Q32 tile can be requested without a rebuild. - scripts/gen_wgsl_headers.py (+ test_wgsl_codegen.py): format-stable generated headers with a regression test. ghstack-source-id: 411961459 @exported-using-ghexport Differential Revision: [D113171744](https://our.internmc.facebook.com/intern/diff/D113171744/)
Pull Request resolved: #21448 **Make WGSL publication rollback-safe without changing generated runtime output** Generated headers and the registry were previously written incrementally, so later validation, interruption, or write failure could leave a partial generated tree. This change stages and validates the complete output set before sequential publication, rejects path/name/C++ symbol collisions, and restores replaced outputs on failure or interruption. The family model follows Vulkan's YAML/template generation. The WebGPU publisher adds staging, backup, rollback, and temporary-file cleanup; it does not claim filesystem-atomic multi-file replacement. Key changes: - `gen_wgsl_headers.py` — stage, validate, publish, restore on failure, and clean temporary outputs. - `test_wgsl_codegen.py` — cover collisions, orphan repair, interruption, rollback, and idempotence. All generated shader bytes and runtime registry entries remain unchanged. Co-authored-with: Claude Code. ghstack-source-id: 411961465 @exported-using-ghexport Differential Revision: [D113979658](https://our.internmc.facebook.com/intern/diff/D113979658/)
Pull Request resolved: #21449 **Generate typed `to_copy` variants from one byte-identical template** The two conversion directions duplicated the same WGSL structure and could drift independently. This moves fp32-to-int32 and int32-to-fp32 into one typed family while preserving both expanded payloads. Mirrors Vulkan `backends/vulkan/runtime/graph/ops/glsl/view_convert_buffer.{glsl,yaml}`. Key changes: - `to_copy_convert.wgsl` and YAML — declare the two typed variants. - Generated headers and codegen locks — preserve symbols, workgroups, registry entries, and payload hashes. - Structural/native tests — lock both directions and the round trip. Host routing and dispatch remain unchanged; tests expand fixture coverage. Co-authored-with: Claude Code. ghstack-source-id: 411961470 @exported-using-ghexport Differential Revision: [D113979712](https://our.internmc.facebook.com/intern/diff/D113979712/)
Pull Request resolved: #21450 **Generate extrema and unary shader families** The extrema reductions and ten no-parameter unary kernels duplicated shader skeletons that could drift independently. This consolidates amax/amin behind one extrema template and abs/cos/exp/hardswish/neg/round/rsqrt/sin/sqrt/tanh behind one unary template while preserving the generated runtime payloads. Key changes: - Generate amax/amin from one extrema manifest. - Generate ten unary payloads from one operator-expression manifest. - Lock expanded bytes, registry entries, delegation, and boundary numerics. The attempted Unary lifecycle migration is intentionally not part of the stack: its performance campaign did not produce an authoritative passing result, so the Unary builder, interface, and activation/sigmoid call sites are restored to their pre-migration bytes. Co-authored-with: Claude Code. ghstack-source-id: 411961475 @exported-using-ghexport Differential Revision: [D113979760](https://our.internmc.facebook.com/intern/diff/D113979760/)
…riants Pull Request resolved: #21451 **Generate byte-identical logical and arithmetic binary variants from shared WGSL families** Logical AND/OR and four arithmetic binary kernels duplicated shader skeletons and broadcast logic. This consolidates logical AND/OR behind one packed-Boolean family and minimum/pow/floor_divide/mul into the existing binary family, with a permanent mixed-rank broadcast contract. Key changes: - Generate logical AND/OR from one operator-token manifest. - Generate minimum, pow, floor_divide, and mul beside the existing div/sub variants. - Lock same-shape and mixed-rank expressions, exact payloads/workgroups, PTE delegation, and broadcast boundary cases. No runtime C++ dispatch, bindings, pipeline construction, workgroups, or expanded shader payloads change. Four standalone WGSL inputs are removed, and future compatible variants require manifest entries instead of copied kernels. This follows the Vulkan binary-family pattern. Co-authored-with: Claude Code. ghstack-source-id: 411961479 @exported-using-ghexport Differential Revision: [D113979789](https://our.internmc.facebook.com/intern/diff/D113979789/)
…ynamic resize Pull Request resolved: #21483 **Preserve immutable Linear and Q4 embedding shader controls across dynamic resize through typed parameter authorities.** **Problem** The Linear resize hook dropped has_bias, and the Q4 embedding resize path could drop is_linear_weight. Both defects produced correct build-time outputs but silently changed semantics after a live shape update. **Solution** - Before: build and resize populated control words independently, allowing immutable shader state to reset. - After: each build/resize pair calls one typed helper while recomputing only live counts and dispatch. **Implementation** - Linear.cpp — make_linear_params preserves the bias flag across vec4 and tiled routes. - EmbeddingQ4gsw.cpp — EmbeddingLayout and make_embedding_params preserve nibble layout across resize. - Mirrors Vulkan runtime/graph/ops/impl/Linear.cpp and EmbeddingQ4gsw.cpp, which carry immutable shader controls through dynamic dispatch. **Constraints** WGSL bytes, shader selection, bindings, uniform sizes, queue-write counts, dispatch formulas, and pipeline topology are unchanged. Production handler code is smaller because duplicated field population and scalar captures are removed. Co-authored-with: Claude Code. ghstack-source-id: 411961489 @exported-using-ghexport Differential Revision: [D113992326](https://our.internmc.facebook.com/intern/diff/D113992326/)
This PR was created by the merge bot to help merge the original PR into the main branch. ghstack PR number: #21597 by @JCNTH ^ Please use this as the source of truth for the PR details, comments, and reviews ghstack PR base: https://github.com/pytorch/executorch/tree/gh/JCNTH/204/base ghstack PR head: https://github.com/pytorch/executorch/tree/gh/JCNTH/204/head Merge bot PR base: https://github.com/pytorch/executorch/tree/gh/JCNTH/203/orig Merge bot PR head: https://github.com/pytorch/executorch/tree/gh/JCNTH/204/orig Differential Revision: [D114936148](https://our.internmc.facebook.com/intern/diff/D114936148/) @diff-train-skip-merge cc @SS-JIA @manuelcandales @digantdesai @cbilgin --------- Co-authored-by: Julian Ng-Thow-Hing <juliannth@meta.com> Co-authored-by: Julian Ng-Thow-Hing <107437036+JCNTH@users.noreply.github.com>
JCNTH
requested review from
SS-JIA,
kirklandsign and
larryliu0820
as code owners
August 7, 2026 17:50
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
This PR was created by the merge bot to help merge the original PR into the main branch.
ghstack PR number: #21130 by @JCNTH
^ Please use this as the source of truth for the PR details, comments, and reviews
ghstack PR base: https://github.com/pytorch/executorch/tree/gh/JCNTH/106/base
ghstack PR head: https://github.com/pytorch/executorch/tree/gh/JCNTH/106/head
Merge bot PR base: https://github.com/pytorch/executorch/tree/gh/JCNTH/105/orig
Merge bot PR head: https://github.com/pytorch/executorch/tree/gh/JCNTH/106/orig
@diff-train-skip-merge