Skip to content

[ExecuTorch][WebGPU] Add certified scoped output suppression - #21646

Merged
JCNTH merged 17 commits into
mainfrom
gh/JCNTH/103/orig
Aug 7, 2026
Merged

[ExecuTorch][WebGPU] Add certified scoped output suppression#21646
JCNTH merged 17 commits into
mainfrom
gh/JCNTH/103/orig

Conversation

@pytorchbot

Copy link
Copy Markdown
Collaborator

This PR was created by the merge bot to help merge the original PR into the main branch.
ghstack PR number: #21127 by @JCNTH
^ Please use this as the source of truth for the PR details, comments, and reviews
ghstack PR base: https://github.com/pytorch/executorch/tree/gh/JCNTH/103/base
ghstack PR head: https://github.com/pytorch/executorch/tree/gh/JCNTH/103/head
Merge bot PR base: https://github.com/pytorch/executorch/tree/main
Merge bot PR head: https://github.com/pytorch/executorch/tree/gh/JCNTH/103/orig

@diff-train-skip-merge

Pull Request resolved: #21127

When a caller discards a model's terminal output tensor, the backend still
dispatches the final quantized-linear range and reads it back, wasting a
dispatch and a device-to-host copy on every execution. This adds a default-off,
certificate-bound execution option that suppresses exactly one explicitly
certified terminal output dispatch range and its readback. The certificate must
bind the exact PTE and method and prove a single delegate, no portable nodes,
and a unique leaf output resolving to the caller's host buffer; anything
ambiguous, aliased, portable, or uncertified fails closed and executes
normally. No Vulkan analogue (WebGPU runtime output-contract change).

Key changes:
- WebGPUExecutionOptions.{h,cpp}: certificate struct, RAII
  ScopedWebGPUExecutionOptions, and plan_webgpu_execution suppression planner.
- WebGPUGraph.{cpp,h}: resolve the certified output ordinal and skip its
  dispatch range and readback when suppression is active.
- WebGPUBackend.cpp: thread the scoped option through execution.
ghstack-source-id: 411961424
@exported-using-ghexport

Differential Revision: [D113171751](https://our.internmc.facebook.com/intern/diff/D113171751/)
@pytorch-bot

pytorch-bot Bot commented Aug 7, 2026

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/21646

Note: Links to docs will display an error until the docs builds have been completed.

⏳ No Failures, 147 Pending

As of commit 4706240 with merge base 9bff727 (image):
💚 Looks good so far! There are no failures yet. 💚

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@github-actions

github-actions Bot commented Aug 7, 2026

Copy link
Copy Markdown

This PR needs a release notes: label

If your change should be included in the release notes (i.e. would users of this library care about this change?), please use a label starting with release notes:. This helps us keep track and include your important work in the next release notes.

To add a label, you can comment to pytorchbot, for example
@pytorchbot label "release notes: none"

For more information, see
https://github.com/pytorch/pytorch/wiki/PyTorch-AutoLabel-Bot#why-categorize-for-release-notes-and-how-does-it-work.

@meta-cla meta-cla Bot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Aug 7, 2026
JCNTH and others added 15 commits August 7, 2026 10:51
Pull Request resolved: #21128

**Centralize dynamic dispatch routing, resize hooks, and workgroup-grid ownership**

Quantized-linear and attention optimizations select specialized dispatch ranges from live sequence lengths, but the original handlers duplicated resize bookkeeping and exposed dispatch indices. This change centralizes route coordination and dynamic-grid ownership in `WebGPUGraph`.

Key changes:
- `WebGPUUtils.h` — register mutually exclusive routes and validate their ranges.
- `WebGPUGraph.{h,cpp}` — add typed named resize hooks, descriptor-driven compute construction, graph-owned parameter buffers and dynamic grids, staged grid updates, and exception-safe retry.
- `QuantizedLinear.cpp`, `Sdpa.cpp`, and `SdpaFdDecode.{h,cpp}` — select specialized ranges through the shared route registry.
- `test_compute_dispatch.cpp` — cover invalid triggers, route overlap, rollback/retry, ownership, and repeated transitions.

The route registry is WebGPU-specific. Graph-owned grid recomputation mirrors Vulkan's `DynamicDispatchNode` lifecycle without introducing a Vulkan-style node hierarchy.

Co-authored-with: Claude Code.
ghstack-source-id: 411961429
@exported-using-ghexport

Differential Revision: [D113171748](https://our.internmc.facebook.com/intern/diff/D113171748/)
Pull Request resolved: #21129

The SwiGLU feed-forward activation lowered to separate gate and up projections
followed by a sigmoid and two elementwise multiplies, launching several
dispatches over the same tensors. This fuses the sigmoid and the two multiplies
into a single dynamically-resized 2D dispatch that computes
(gate * sigmoid(gate)) * up in one pass. Fusion applies only when the pattern
is exact and its scratch is planner-owned; graph outputs, extra consumers, and
separate inputs fall back to the unfused chain, and the landed broad-QKV
behavior is left untouched until its dedicated corrective diff. No Vulkan
analogue: the Vulkan backend has no fused SiLU/SwiGLU shader and composes it
from a sigmoid unary op plus an elementwise multiply at graph level.

Key changes:
- runtime/ops/mul/silu_mul_fused.wgsl (+ generated header): one-thread-per-
  element SiLU-gated multiply of gate and up into a single output.
- WebGPUGraph.cpp: detect the exact gate/up/sigmoid/mul pattern, own the gate
  scratch, size the 2D dispatch, and fall back conservatively.
ghstack-source-id: 411961441
@exported-using-ghexport

Differential Revision: [D113171742](https://our.internmc.facebook.com/intern/diff/D113171742/)
Pull Request resolved: #21130

Llama prefill drives the ordinary quantized-linear projections at a small set
of fixed batch-row counts, and the generic Steel schedule leaves throughput on
the table for them. This adds a frozen BK64 Steel quantized-GEMM route for the
exact accepted Llama ordinary-projection shapes at live M128, M508, and M512,
selected only when the capability and dynamic-route guards all pass. M511 and
other prefill sizes stay on the generic Steel schedule and M1 stays on bicol
decode, so the route fails closed outside its accepted shapes. Mirrors Vulkan
xplat/executorch/backends/vulkan/runtime/graph/ops/glsl/q4gsw_linear_gemm__w_4x8.glsl
(the 4-bit grouped-symmetric-weight GEMM), specialized here for the BK64 tile.

Key changes:
- runtime/ops/quantized_linear/q4gsw_steel_bk64.wgsl (+ generated header): the
  BK64-tiled q4gsw prefill GEMM kernel.
- QuantizedLinear.cpp, WebGPUUtils.h: exact-shape predicate, capability guard,
  and dynamic re-entry into and out of the BK64 route.
ghstack-source-id: 411961445
@exported-using-ghexport

Differential Revision: [D113171739](https://our.internmc.facebook.com/intern/diff/D113171739/)
Pull Request resolved: #21131

The previously landed broad QKV fusion applied too widely and did not match the
BK64 schedule now used for the ordinary projections. This corrective diff
replaces it with a capability- and geometry-qualified BK64 kernel that fuses the
exact Llama Q/K/V projection triple at live M128, M508, and M512, packing the
constant weights and scales once and scattering the result into three distinct
planner-safe outputs. Outside the accepted shapes it switches atomically back
to the ordinary Steel and bicol routes, and it deletes the obsolete broad
shader and header so only one QKV path remains. Mirrors Vulkan
xplat/executorch/backends/vulkan/runtime/graph/ops/glsl/q4gsw_linear_gemm__w_4x8.glsl
for the per-projection GEMM; the three-output fusion itself is WebGPU-specific.

Key changes:
- Adds runtime/ops/quantized_linear/q4gsw_qkv_bk64.wgsl (+ generated header);
  removes the obsolete q4gsw_linear_gemm_qkv_fused shader and header.
- WebGPUGraph.{cpp,h}: geometry and capability predicate, one-time constant
  packing, distinct Q/K/V outputs, and atomic fallback to Steel/bicol.
ghstack-source-id: 411961443
@exported-using-ghexport

Differential Revision: [D113171749](https://our.internmc.facebook.com/intern/diff/D113171749/)
Pull Request resolved: #21132

Materializing the full attention-weight matrix for Llama prefill is memory- and
bandwidth-heavy and does not scale to longer sequences. This adds a single-pass
online-softmax K16 causal-attention kernel for the exact Llama Hq32/Hkv8/G4/D64
geometry with fp16 KV, guarded by explicit adapter limits, so accepted prefill
shapes through S512 never materialize attention weights. S1 keeps the existing
FlashDecoding path, and unsupported geometry, storage, or capabilities fall back
to the materialized implementation. No Vulkan analogue (WebGPU-specific): the
Vulkan backend has only a materialized attention (compute the weights, then a
separate multi-pass softmax), with no online-softmax kernel.

Key changes:
- runtime/ops/sdpa/streaming_attention_k16_causal_bound.wgsl (+ generated
  header): the online-softmax single-pass causal kernel.
- Sdpa.cpp: exact Llama geometry and capability guards, S512 transition
  handling, and fallback to the materialized route.
ghstack-source-id: 411961446
@exported-using-ghexport

Differential Revision: [D113171738](https://our.internmc.facebook.com/intern/diff/D113171738/)
Pull Request resolved: #21133

Building the WebGPU backend for the browser previously required a manual
source-tree patch to pick the right WebGPU implementation. This makes the CMake
configuration select the emdawnwebgpu port for Emscripten builds while native
builds keep Dawn and their platform libraries, so the production kernel stack
builds for the browser unpatched. Build-system integration only; no model or
kernel behavior changes.

Key changes:
- CMakeLists.txt: choose emdawnwebgpu under Emscripten, Dawn otherwise.
- test/test_cmake_configuration.py: contract test asserting the selection.
ghstack-source-id: 411961450
@exported-using-ghexport

Differential Revision: [D113171750](https://our.internmc.facebook.com/intern/diff/D113171750/)
Pull Request resolved: #21134

With dynamic sequence positions, the backend could record only one SDPA route,
so a decode following a dynamic-position prefill could dispatch the wrong
kernel, and there was no way to confirm which kernel actually ran. This treats a
dynamic SymInt position (not just a dynamic sequence dimension) as sufficient to
record both SDPA routes, exposes active-kernel route attestation for correctness
and performance harnesses, and reads the timestamp-query gate per execution so
diagnostics can be enabled after module initialization. The attestation state
and query are compiled only under the WGPU_BACKEND_ENABLE_PROFILING build flag,
so production builds carry no additional state or cost, mirroring how the Vulkan
backend gates its QueryPool behind ET_EVENT_TRACER_ENABLED. No Vulkan analogue
for the routing fix (WebGPU runtime routing and observability).

Key changes:
- WebGPUGraph.{cpp,h}: expose the active-route attestation query (compile-gated
  behind WGPU_BACKEND_ENABLE_PROFILING) and read the timestamp-query gate per
  execution.
- WebGPUUtils.h, Sdpa.cpp: treat a dynamic SymInt position (not only a dynamic
  sequence dimension) as sufficient to record both SDPA routes.
ghstack-source-id: 411961454
@exported-using-ghexport

Differential Revision: [D113171745](https://our.internmc.facebook.com/intern/diff/D113171745/)
Pull Request resolved: #21135

**Add HuggingFace rotate-half RoPE and migrate RoPE to shared dispatch construction**

Qwen models pair the first and second halves of each head vector rather than adjacent elements. This adds the rotate-half operator with dynamic start-position updates, full-dimension Q/K handling, generated WGSL, and strict frequency-table bounds.

The shared RoPE handler also closes the dispatch-boilerplate review: it uses `graph.device()`, typed graph-owned parameter buffers, descriptor-driven bindings and pipelines, named validation and resize callbacks, generated shader-registry lookup, scoped ownership, and graph-owned workgroup recomputation. Interleaved shader math and output are unchanged.

Key changes:
- `rotary_embedding_hf.wgsl` and generated registry entry — one thread per pair for HuggingFace rotate-half.
- `RotaryEmbedding.cpp` — named validation, typed resize contexts, shared dispatch descriptors, and one initial/resize grid picker per route.
- Native tests — malformed input rejection, dynamic bounds, shrink/regrow reuse, and HF/interleaved lifecycle coverage.

Co-authored-with: Claude Code.
ghstack-source-id: 411961455
@exported-using-ghexport

Differential Revision: [D113171746](https://our.internmc.facebook.com/intern/diff/D113171746/)
Pull Request resolved: #21136

Qwen3's attention geometry differs from Llama's, and its KV cache is produced in
fp32 on the host but must be consumed in fp16 on the device. This adds guarded
Qwen3 K16 streaming online-softmax attention schedules — a Q16 schedule that is
the automatic default whenever the exact geometry and capability guards pass,
plus a Q32 candidate that is opt-in through the `sdpa_query_tile` runtime spec
(BackendOption) for future autotuning — together with the exact fp32-host to
fp16-device KV-cache boundary conversion. Selection requires the exact Qwen3
geometry, fp16 KV storage, adapter limits, a valid workgroup count, and an exact
2:1 byte ratio; the established Llama, materialized, and FlashDecoding routes
remain fallbacks. This builds on the HuggingFace rotate-half
RoPE operator. No Vulkan analogue (WebGPU-specific online-softmax attention;
Vulkan has only a materialized attention). It also makes the long generated WGSL
provenance and constant declarations format-stable and covers them with a
generator regression test.

Key changes:
- runtime/ops/sdpa/streaming_attention_qwen3_k16_causal_bound.wgsl and
  streaming_attention_qwen3_q32_k16_causal_bound.wgsl (+ generated headers): the
  Q16 and Q32 online-softmax Qwen3 kernels.
- Sdpa.cpp, WebGPUGraph.{cpp,h}: exact Qwen3 geometry and limit guards, Q16
  default route selection, and the fp32-host to fp16-device KV-cache conversion.
- WebGPUBackend.cpp: read the optional `sdpa_query_tile` runtime spec and thread
  it into graph build so the Q32 tile can be requested without a rebuild.
- scripts/gen_wgsl_headers.py (+ test_wgsl_codegen.py): format-stable generated
  headers with a regression test.
ghstack-source-id: 411961459
@exported-using-ghexport

Differential Revision: [D113171744](https://our.internmc.facebook.com/intern/diff/D113171744/)
Pull Request resolved: #21448

**Make WGSL publication rollback-safe without changing generated runtime output**

Generated headers and the registry were previously written incrementally, so later validation, interruption, or write failure could leave a partial generated tree. This change stages and validates the complete output set before sequential publication, rejects path/name/C++ symbol collisions, and restores replaced outputs on failure or interruption.

The family model follows Vulkan's YAML/template generation. The WebGPU publisher adds staging, backup, rollback, and temporary-file cleanup; it does not claim filesystem-atomic multi-file replacement.

Key changes:
- `gen_wgsl_headers.py` — stage, validate, publish, restore on failure, and clean temporary outputs.
- `test_wgsl_codegen.py` — cover collisions, orphan repair, interruption, rollback, and idempotence.

All generated shader bytes and runtime registry entries remain unchanged.

Co-authored-with: Claude Code.
ghstack-source-id: 411961465
@exported-using-ghexport

Differential Revision: [D113979658](https://our.internmc.facebook.com/intern/diff/D113979658/)
Pull Request resolved: #21449

**Generate typed `to_copy` variants from one byte-identical template**

The two conversion directions duplicated the same WGSL structure and could drift independently. This moves fp32-to-int32 and int32-to-fp32 into one typed family while preserving both expanded payloads. Mirrors Vulkan `backends/vulkan/runtime/graph/ops/glsl/view_convert_buffer.{glsl,yaml}`.

Key changes:
- `to_copy_convert.wgsl` and YAML — declare the two typed variants.
- Generated headers and codegen locks — preserve symbols, workgroups, registry entries, and payload hashes.
- Structural/native tests — lock both directions and the round trip.

Host routing and dispatch remain unchanged; tests expand fixture coverage.

Co-authored-with: Claude Code.
ghstack-source-id: 411961470
@exported-using-ghexport

Differential Revision: [D113979712](https://our.internmc.facebook.com/intern/diff/D113979712/)
Pull Request resolved: #21450

**Generate extrema and unary shader families**

The extrema reductions and ten no-parameter unary kernels duplicated shader skeletons that could drift independently. This consolidates amax/amin behind one extrema template and abs/cos/exp/hardswish/neg/round/rsqrt/sin/sqrt/tanh behind one unary template while preserving the generated runtime payloads.

Key changes:
- Generate amax/amin from one extrema manifest.
- Generate ten unary payloads from one operator-expression manifest.
- Lock expanded bytes, registry entries, delegation, and boundary numerics.

The attempted Unary lifecycle migration is intentionally not part of the stack: its performance campaign did not produce an authoritative passing result, so the Unary builder, interface, and activation/sigmoid call sites are restored to their pre-migration bytes.

Co-authored-with: Claude Code.
ghstack-source-id: 411961475
@exported-using-ghexport

Differential Revision: [D113979760](https://our.internmc.facebook.com/intern/diff/D113979760/)
…riants

Pull Request resolved: #21451

**Generate byte-identical logical and arithmetic binary variants from shared WGSL families**

Logical AND/OR and four arithmetic binary kernels duplicated shader skeletons and broadcast logic. This consolidates logical AND/OR behind one packed-Boolean family and minimum/pow/floor_divide/mul into the existing binary family, with a permanent mixed-rank broadcast contract.

Key changes:
- Generate logical AND/OR from one operator-token manifest.
- Generate minimum, pow, floor_divide, and mul beside the existing div/sub variants.
- Lock same-shape and mixed-rank expressions, exact payloads/workgroups, PTE delegation, and broadcast boundary cases.

No runtime C++ dispatch, bindings, pipeline construction, workgroups, or expanded shader payloads change. Four standalone WGSL inputs are removed, and future compatible variants require manifest entries instead of copied kernels. This follows the Vulkan binary-family pattern.

Co-authored-with: Claude Code.
ghstack-source-id: 411961479
@exported-using-ghexport

Differential Revision: [D113979789](https://our.internmc.facebook.com/intern/diff/D113979789/)
…ynamic resize

Pull Request resolved: #21483

**Preserve immutable Linear and Q4 embedding shader controls across dynamic resize through typed parameter authorities.**

**Problem**
The Linear resize hook dropped has_bias, and the Q4 embedding resize path could drop is_linear_weight. Both defects produced correct build-time outputs but silently changed semantics after a live shape update.

**Solution**
- Before: build and resize populated control words independently, allowing immutable shader state to reset.
- After: each build/resize pair calls one typed helper while recomputing only live counts and dispatch.

**Implementation**
- Linear.cpp — make_linear_params preserves the bias flag across vec4 and tiled routes.
- EmbeddingQ4gsw.cpp — EmbeddingLayout and make_embedding_params preserve nibble layout across resize.
- Mirrors Vulkan runtime/graph/ops/impl/Linear.cpp and EmbeddingQ4gsw.cpp, which carry immutable shader controls through dynamic dispatch.

**Constraints**
WGSL bytes, shader selection, bindings, uniform sizes, queue-write counts, dispatch formulas, and pipeline topology are unchanged. Production handler code is smaller because duplicated field population and scalar captures are removed.

Co-authored-with: Claude Code.
ghstack-source-id: 411961489
@exported-using-ghexport

Differential Revision: [D113992326](https://our.internmc.facebook.com/intern/diff/D113992326/)
This PR was created by the merge bot to help merge the original PR into
the main branch.
ghstack PR number: #21597 by
@JCNTH
^ Please use this as the source of truth for the PR details, comments,
and reviews
ghstack PR base:
https://github.com/pytorch/executorch/tree/gh/JCNTH/204/base
ghstack PR head:
https://github.com/pytorch/executorch/tree/gh/JCNTH/204/head
Merge bot PR base:
https://github.com/pytorch/executorch/tree/gh/JCNTH/203/orig
Merge bot PR head:
https://github.com/pytorch/executorch/tree/gh/JCNTH/204/orig
Differential Revision:
[D114936148](https://our.internmc.facebook.com/intern/diff/D114936148/)
@diff-train-skip-merge

cc @SS-JIA @manuelcandales @digantdesai @cbilgin

---------

Co-authored-by: Julian Ng-Thow-Hing <juliannth@meta.com>
Co-authored-by: Julian Ng-Thow-Hing <107437036+JCNTH@users.noreply.github.com>
@JCNTH
JCNTH self-requested a review August 7, 2026 17:51
@JCNTH
JCNTH merged commit ceca90f into main Aug 7, 2026
179 of 181 checks passed
@JCNTH
JCNTH deleted the gh/JCNTH/103/orig branch August 7, 2026 17:58
JCNTH added a commit that referenced this pull request Aug 7, 2026
The squash merge for #21646 duplicated the fp16 upload and BOOL
conversion paths and dropped three native route tests. Restore the
source that matches the reviewed fbsource stack, including
compile-gating timestamp profiling.

Test Plan: exact tree comparison against the reconstructed 21-diff
stack; git diff --check; GitHub CI.

Co-authored-with: OpenAI Codex.

Signed-off-by: Julian Ng-Thow-Hing <juliannth@meta.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants