Skip to content

[ExecuTorch][WebGPU] Repair dynamic SDPA routing, add attestation - #21134

Merged
meta-codesync[bot] merged 11 commits into
gh/JCNTH/110/basefrom
gh/JCNTH/110/head
Aug 7, 2026
Merged

[ExecuTorch][WebGPU] Repair dynamic SDPA routing, add attestation#21134
meta-codesync[bot] merged 11 commits into
gh/JCNTH/110/basefrom
gh/JCNTH/110/head

Conversation

@JCNTH

@JCNTH JCNTH commented Jul 22, 2026

Copy link
Copy Markdown
Contributor

Stack from ghstack (oldest at bottom):

With dynamic sequence positions, the backend could record only one SDPA route,
so a decode following a dynamic-position prefill could dispatch the wrong
kernel, and there was no way to confirm which kernel actually ran. This treats a
dynamic SymInt position (not just a dynamic sequence dimension) as sufficient to
record both SDPA routes, exposes active-kernel route attestation for correctness
and performance harnesses, and reads the timestamp-query gate per execution so
diagnostics can be enabled after module initialization. The attestation state
and query are compiled only under the WGPU_BACKEND_ENABLE_PROFILING build flag,
so production builds carry no additional state or cost, mirroring how the Vulkan
backend gates its QueryPool behind ET_EVENT_TRACER_ENABLED. No Vulkan analogue
for the routing fix (WebGPU runtime routing and observability).

Key changes:

  • WebGPUGraph.{cpp,h}: expose the active-route attestation query (compile-gated
    behind WGPU_BACKEND_ENABLE_PROFILING) and read the timestamp-query gate per
    execution.
  • WebGPUUtils.h, Sdpa.cpp: treat a dynamic SymInt position (not only a dynamic
    sequence dimension) as sufficient to record both SDPA routes.
    @exported-using-ghexport

Differential Revision: D113171745

Differential Revision: D113171745

[ghstack-poisoned]
@pytorch-bot

pytorch-bot Bot commented Jul 22, 2026

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/21134

Note: Links to docs will display an error until the docs builds have been completed.

❌ 2 New Failures, 35 Pending

As of commit 3d8e640 with merge base 28a7fac (image):

NEW FAILURES - The following jobs have failed:

  • Cadence Build & Test / cpu-test / test-aot / test-aot (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID AAAC:330715:357FCBE:B6F4ACF:6A760B4E and timestamp 2026-08-07 16:43:58 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • Cadence Build & Test / cpu-test / test-ops / test-ops (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID 9CE0:1600F6:37632DB:BC43A60:6A760B4F and timestamp 2026-08-07 16:43:59 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@github-actions

Copy link
Copy Markdown

This PR needs a release notes: label

If your change should be included in the release notes (i.e. would users of this library care about this change?), please use a label starting with release notes:. This helps us keep track and include your important work in the next release notes.

To add a label, you can comment to pytorchbot, for example
@pytorchbot label "release notes: none"

For more information, see
https://github.com/pytorch/pytorch/wiki/PyTorch-AutoLabel-Bot#why-categorize-for-release-notes-and-how-does-it-work.

@meta-cla meta-cla Bot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Jul 22, 2026
[ghstack-poisoned]
[ghstack-poisoned]
[ghstack-poisoned]

@SS-JIA SS-JIA left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review automatically exported from Phabricator review in Meta.

[ghstack-poisoned]
[ghstack-poisoned]
[ghstack-poisoned]
  fbsource master

[ghstack-poisoned]
@meta-codesync
meta-codesync Bot merged commit 8503b32 into gh/JCNTH/110/base Aug 7, 2026
181 of 183 checks passed
@meta-codesync
meta-codesync Bot deleted the gh/JCNTH/110/head branch August 7, 2026 17:20
@meta-codesync
meta-codesync Bot temporarily deployed to cherry-pick-bot August 7, 2026 17:20 Inactive
JCNTH added a commit that referenced this pull request Aug 7, 2026
Pull Request resolved: #21134

With dynamic sequence positions, the backend could record only one SDPA route,
so a decode following a dynamic-position prefill could dispatch the wrong
kernel, and there was no way to confirm which kernel actually ran. This treats a
dynamic SymInt position (not just a dynamic sequence dimension) as sufficient to
record both SDPA routes, exposes active-kernel route attestation for correctness
and performance harnesses, and reads the timestamp-query gate per execution so
diagnostics can be enabled after module initialization. The attestation state
and query are compiled only under the WGPU_BACKEND_ENABLE_PROFILING build flag,
so production builds carry no additional state or cost, mirroring how the Vulkan
backend gates its QueryPool behind ET_EVENT_TRACER_ENABLED. No Vulkan analogue
for the routing fix (WebGPU runtime routing and observability).

Key changes:
- WebGPUGraph.{cpp,h}: expose the active-route attestation query (compile-gated
  behind WGPU_BACKEND_ENABLE_PROFILING) and read the timestamp-query gate per
  execution.
- WebGPUUtils.h, Sdpa.cpp: treat a dynamic SymInt position (not only a dynamic
  sequence dimension) as sufficient to record both SDPA routes.
ghstack-source-id: 411961454
@exported-using-ghexport

Differential Revision: [D113171745](https://our.internmc.facebook.com/intern/diff/D113171745/)
JCNTH added a commit that referenced this pull request Aug 7, 2026
Pull Request resolved: #21134

With dynamic sequence positions, the backend could record only one SDPA route,
so a decode following a dynamic-position prefill could dispatch the wrong
kernel, and there was no way to confirm which kernel actually ran. This treats a
dynamic SymInt position (not just a dynamic sequence dimension) as sufficient to
record both SDPA routes, exposes active-kernel route attestation for correctness
and performance harnesses, and reads the timestamp-query gate per execution so
diagnostics can be enabled after module initialization. The attestation state
and query are compiled only under the WGPU_BACKEND_ENABLE_PROFILING build flag,
so production builds carry no additional state or cost, mirroring how the Vulkan
backend gates its QueryPool behind ET_EVENT_TRACER_ENABLED. No Vulkan analogue
for the routing fix (WebGPU runtime routing and observability).

Key changes:
- WebGPUGraph.{cpp,h}: expose the active-route attestation query (compile-gated
  behind WGPU_BACKEND_ENABLE_PROFILING) and read the timestamp-query gate per
  execution.
- WebGPUUtils.h, Sdpa.cpp: treat a dynamic SymInt position (not only a dynamic
  sequence dimension) as sufficient to record both SDPA routes.
ghstack-source-id: 411961454
@exported-using-ghexport

Differential Revision: [D113171745](https://our.internmc.facebook.com/intern/diff/D113171745/)
JCNTH added a commit that referenced this pull request Aug 7, 2026
Pull Request resolved: #21134

With dynamic sequence positions, the backend could record only one SDPA route,
so a decode following a dynamic-position prefill could dispatch the wrong
kernel, and there was no way to confirm which kernel actually ran. This treats a
dynamic SymInt position (not just a dynamic sequence dimension) as sufficient to
record both SDPA routes, exposes active-kernel route attestation for correctness
and performance harnesses, and reads the timestamp-query gate per execution so
diagnostics can be enabled after module initialization. The attestation state
and query are compiled only under the WGPU_BACKEND_ENABLE_PROFILING build flag,
so production builds carry no additional state or cost, mirroring how the Vulkan
backend gates its QueryPool behind ET_EVENT_TRACER_ENABLED. No Vulkan analogue
for the routing fix (WebGPU runtime routing and observability).

Key changes:
- WebGPUGraph.{cpp,h}: expose the active-route attestation query (compile-gated
  behind WGPU_BACKEND_ENABLE_PROFILING) and read the timestamp-query gate per
  execution.
- WebGPUUtils.h, Sdpa.cpp: treat a dynamic SymInt position (not only a dynamic
  sequence dimension) as sufficient to record both SDPA routes.
ghstack-source-id: 411961454
@exported-using-ghexport

Differential Revision: [D113171745](https://our.internmc.facebook.com/intern/diff/D113171745/)
JCNTH added a commit that referenced this pull request Aug 7, 2026
Pull Request resolved: #21134

With dynamic sequence positions, the backend could record only one SDPA route,
so a decode following a dynamic-position prefill could dispatch the wrong
kernel, and there was no way to confirm which kernel actually ran. This treats a
dynamic SymInt position (not just a dynamic sequence dimension) as sufficient to
record both SDPA routes, exposes active-kernel route attestation for correctness
and performance harnesses, and reads the timestamp-query gate per execution so
diagnostics can be enabled after module initialization. The attestation state
and query are compiled only under the WGPU_BACKEND_ENABLE_PROFILING build flag,
so production builds carry no additional state or cost, mirroring how the Vulkan
backend gates its QueryPool behind ET_EVENT_TRACER_ENABLED. No Vulkan analogue
for the routing fix (WebGPU runtime routing and observability).

Key changes:
- WebGPUGraph.{cpp,h}: expose the active-route attestation query (compile-gated
  behind WGPU_BACKEND_ENABLE_PROFILING) and read the timestamp-query gate per
  execution.
- WebGPUUtils.h, Sdpa.cpp: treat a dynamic SymInt position (not only a dynamic
  sequence dimension) as sufficient to record both SDPA routes.
ghstack-source-id: 411961454
@exported-using-ghexport

Differential Revision: [D113171745](https://our.internmc.facebook.com/intern/diff/D113171745/)
JCNTH added a commit that referenced this pull request Aug 7, 2026
Pull Request resolved: #21134

With dynamic sequence positions, the backend could record only one SDPA route,
so a decode following a dynamic-position prefill could dispatch the wrong
kernel, and there was no way to confirm which kernel actually ran. This treats a
dynamic SymInt position (not just a dynamic sequence dimension) as sufficient to
record both SDPA routes, exposes active-kernel route attestation for correctness
and performance harnesses, and reads the timestamp-query gate per execution so
diagnostics can be enabled after module initialization. The attestation state
and query are compiled only under the WGPU_BACKEND_ENABLE_PROFILING build flag,
so production builds carry no additional state or cost, mirroring how the Vulkan
backend gates its QueryPool behind ET_EVENT_TRACER_ENABLED. No Vulkan analogue
for the routing fix (WebGPU runtime routing and observability).

Key changes:
- WebGPUGraph.{cpp,h}: expose the active-route attestation query (compile-gated
  behind WGPU_BACKEND_ENABLE_PROFILING) and read the timestamp-query gate per
  execution.
- WebGPUUtils.h, Sdpa.cpp: treat a dynamic SymInt position (not only a dynamic
  sequence dimension) as sufficient to record both SDPA routes.
ghstack-source-id: 411961454
@exported-using-ghexport

Differential Revision: [D113171745](https://our.internmc.facebook.com/intern/diff/D113171745/)
JCNTH added a commit that referenced this pull request Aug 7, 2026
Pull Request resolved: #21134

With dynamic sequence positions, the backend could record only one SDPA route,
so a decode following a dynamic-position prefill could dispatch the wrong
kernel, and there was no way to confirm which kernel actually ran. This treats a
dynamic SymInt position (not just a dynamic sequence dimension) as sufficient to
record both SDPA routes, exposes active-kernel route attestation for correctness
and performance harnesses, and reads the timestamp-query gate per execution so
diagnostics can be enabled after module initialization. The attestation state
and query are compiled only under the WGPU_BACKEND_ENABLE_PROFILING build flag,
so production builds carry no additional state or cost, mirroring how the Vulkan
backend gates its QueryPool behind ET_EVENT_TRACER_ENABLED. No Vulkan analogue
for the routing fix (WebGPU runtime routing and observability).

Key changes:
- WebGPUGraph.{cpp,h}: expose the active-route attestation query (compile-gated
  behind WGPU_BACKEND_ENABLE_PROFILING) and read the timestamp-query gate per
  execution.
- WebGPUUtils.h, Sdpa.cpp: treat a dynamic SymInt position (not only a dynamic
  sequence dimension) as sufficient to record both SDPA routes.
ghstack-source-id: 411961454
@exported-using-ghexport

Differential Revision: [D113171745](https://our.internmc.facebook.com/intern/diff/D113171745/)
JCNTH added a commit that referenced this pull request Aug 7, 2026
Pull Request resolved: #21134

With dynamic sequence positions, the backend could record only one SDPA route,
so a decode following a dynamic-position prefill could dispatch the wrong
kernel, and there was no way to confirm which kernel actually ran. This treats a
dynamic SymInt position (not just a dynamic sequence dimension) as sufficient to
record both SDPA routes, exposes active-kernel route attestation for correctness
and performance harnesses, and reads the timestamp-query gate per execution so
diagnostics can be enabled after module initialization. The attestation state
and query are compiled only under the WGPU_BACKEND_ENABLE_PROFILING build flag,
so production builds carry no additional state or cost, mirroring how the Vulkan
backend gates its QueryPool behind ET_EVENT_TRACER_ENABLED. No Vulkan analogue
for the routing fix (WebGPU runtime routing and observability).

Key changes:
- WebGPUGraph.{cpp,h}: expose the active-route attestation query (compile-gated
  behind WGPU_BACKEND_ENABLE_PROFILING) and read the timestamp-query gate per
  execution.
- WebGPUUtils.h, Sdpa.cpp: treat a dynamic SymInt position (not only a dynamic
  sequence dimension) as sufficient to record both SDPA routes.
ghstack-source-id: 411961454
@exported-using-ghexport

Differential Revision: [D113171745](https://our.internmc.facebook.com/intern/diff/D113171745/)
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. meta-exported

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants