Skip to content

[ExecuTorch][WebGPU] Add HuggingFace rotate-half RoPE - #21135

Merged
meta-codesync[bot] merged 11 commits into
gh/JCNTH/111/basefrom
gh/JCNTH/111/head
Aug 7, 2026
Merged

[ExecuTorch][WebGPU] Add HuggingFace rotate-half RoPE#21135
meta-codesync[bot] merged 11 commits into
gh/JCNTH/111/basefrom
gh/JCNTH/111/head

Conversation

@JCNTH

@JCNTH JCNTH commented Jul 22, 2026

Copy link
Copy Markdown
Contributor

Stack from ghstack (oldest at bottom):

Add HuggingFace rotate-half RoPE and migrate RoPE to shared dispatch construction

Qwen models pair the first and second halves of each head vector rather than adjacent elements. This adds the rotate-half operator with dynamic start-position updates, full-dimension Q/K handling, generated WGSL, and strict frequency-table bounds.

The shared RoPE handler also closes the dispatch-boilerplate review: it uses graph.device(), typed graph-owned parameter buffers, descriptor-driven bindings and pipelines, named validation and resize callbacks, generated shader-registry lookup, scoped ownership, and graph-owned workgroup recomputation. Interleaved shader math and output are unchanged.

Key changes:

  • rotary_embedding_hf.wgsl and generated registry entry — one thread per pair for HuggingFace rotate-half.
  • RotaryEmbedding.cpp — named validation, typed resize contexts, shared dispatch descriptors, and one initial/resize grid picker per route.
  • Native tests — malformed input rejection, dynamic bounds, shrink/regrow reuse, and HF/interleaved lifecycle coverage.

Co-authored-with: Claude Code.
@exported-using-ghexport

Differential Revision: D113171746

Differential Revision: D113171746

[ghstack-poisoned]
@pytorch-bot

pytorch-bot Bot commented Jul 22, 2026

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/21135

Note: Links to docs will display an error until the docs builds have been completed.

❌ 2 New Failures, 29 Pending

As of commit 2372deb with merge base 28a7fac (image):

NEW FAILURES - The following jobs have failed:

  • Cadence Build & Test / cpu-test / test-aot / test-aot (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID A6A0:E0E81:3793C26:BCF63CC:6A760B4A and timestamp 2026-08-07 16:43:54 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • Cadence Build & Test / cpu-test / test-ops / test-ops (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID ECAC:30BA9A:3991E73:C496DF7:6A760B49 and timestamp 2026-08-07 16:43:53 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@github-actions

Copy link
Copy Markdown

This PR needs a release notes: label

If your change should be included in the release notes (i.e. would users of this library care about this change?), please use a label starting with release notes:. This helps us keep track and include your important work in the next release notes.

To add a label, you can comment to pytorchbot, for example
@pytorchbot label "release notes: none"

For more information, see
https://github.com/pytorch/pytorch/wiki/PyTorch-AutoLabel-Bot#why-categorize-for-release-notes-and-how-does-it-work.

@meta-cla meta-cla Bot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Jul 22, 2026
[ghstack-poisoned]
[ghstack-poisoned]
[ghstack-poisoned]

@SS-JIA SS-JIA left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review automatically exported from Phabricator review in Meta.

[ghstack-poisoned]
[ghstack-poisoned]
[ghstack-poisoned]
  fbsource master

[ghstack-poisoned]
@meta-codesync
meta-codesync Bot merged commit 956057c into gh/JCNTH/111/base Aug 7, 2026
180 of 183 checks passed
@meta-codesync
meta-codesync Bot deleted the gh/JCNTH/111/head branch August 7, 2026 17:21
@meta-codesync
meta-codesync Bot temporarily deployed to cherry-pick-bot August 7, 2026 17:21 Inactive
JCNTH added a commit that referenced this pull request Aug 7, 2026
Pull Request resolved: #21135

**Add HuggingFace rotate-half RoPE and migrate RoPE to shared dispatch construction**

Qwen models pair the first and second halves of each head vector rather than adjacent elements. This adds the rotate-half operator with dynamic start-position updates, full-dimension Q/K handling, generated WGSL, and strict frequency-table bounds.

The shared RoPE handler also closes the dispatch-boilerplate review: it uses `graph.device()`, typed graph-owned parameter buffers, descriptor-driven bindings and pipelines, named validation and resize callbacks, generated shader-registry lookup, scoped ownership, and graph-owned workgroup recomputation. Interleaved shader math and output are unchanged.

Key changes:
- `rotary_embedding_hf.wgsl` and generated registry entry — one thread per pair for HuggingFace rotate-half.
- `RotaryEmbedding.cpp` — named validation, typed resize contexts, shared dispatch descriptors, and one initial/resize grid picker per route.
- Native tests — malformed input rejection, dynamic bounds, shrink/regrow reuse, and HF/interleaved lifecycle coverage.

Co-authored-with: Claude Code.
ghstack-source-id: 411961455
@exported-using-ghexport

Differential Revision: [D113171746](https://our.internmc.facebook.com/intern/diff/D113171746/)
JCNTH added a commit that referenced this pull request Aug 7, 2026
Pull Request resolved: #21135

**Add HuggingFace rotate-half RoPE and migrate RoPE to shared dispatch construction**

Qwen models pair the first and second halves of each head vector rather than adjacent elements. This adds the rotate-half operator with dynamic start-position updates, full-dimension Q/K handling, generated WGSL, and strict frequency-table bounds.

The shared RoPE handler also closes the dispatch-boilerplate review: it uses `graph.device()`, typed graph-owned parameter buffers, descriptor-driven bindings and pipelines, named validation and resize callbacks, generated shader-registry lookup, scoped ownership, and graph-owned workgroup recomputation. Interleaved shader math and output are unchanged.

Key changes:
- `rotary_embedding_hf.wgsl` and generated registry entry — one thread per pair for HuggingFace rotate-half.
- `RotaryEmbedding.cpp` — named validation, typed resize contexts, shared dispatch descriptors, and one initial/resize grid picker per route.
- Native tests — malformed input rejection, dynamic bounds, shrink/regrow reuse, and HF/interleaved lifecycle coverage.

Co-authored-with: Claude Code.
ghstack-source-id: 411961455
@exported-using-ghexport

Differential Revision: [D113171746](https://our.internmc.facebook.com/intern/diff/D113171746/)
JCNTH added a commit that referenced this pull request Aug 7, 2026
Pull Request resolved: #21135

**Add HuggingFace rotate-half RoPE and migrate RoPE to shared dispatch construction**

Qwen models pair the first and second halves of each head vector rather than adjacent elements. This adds the rotate-half operator with dynamic start-position updates, full-dimension Q/K handling, generated WGSL, and strict frequency-table bounds.

The shared RoPE handler also closes the dispatch-boilerplate review: it uses `graph.device()`, typed graph-owned parameter buffers, descriptor-driven bindings and pipelines, named validation and resize callbacks, generated shader-registry lookup, scoped ownership, and graph-owned workgroup recomputation. Interleaved shader math and output are unchanged.

Key changes:
- `rotary_embedding_hf.wgsl` and generated registry entry — one thread per pair for HuggingFace rotate-half.
- `RotaryEmbedding.cpp` — named validation, typed resize contexts, shared dispatch descriptors, and one initial/resize grid picker per route.
- Native tests — malformed input rejection, dynamic bounds, shrink/regrow reuse, and HF/interleaved lifecycle coverage.

Co-authored-with: Claude Code.
ghstack-source-id: 411961455
@exported-using-ghexport

Differential Revision: [D113171746](https://our.internmc.facebook.com/intern/diff/D113171746/)
JCNTH added a commit that referenced this pull request Aug 7, 2026
Pull Request resolved: #21135

**Add HuggingFace rotate-half RoPE and migrate RoPE to shared dispatch construction**

Qwen models pair the first and second halves of each head vector rather than adjacent elements. This adds the rotate-half operator with dynamic start-position updates, full-dimension Q/K handling, generated WGSL, and strict frequency-table bounds.

The shared RoPE handler also closes the dispatch-boilerplate review: it uses `graph.device()`, typed graph-owned parameter buffers, descriptor-driven bindings and pipelines, named validation and resize callbacks, generated shader-registry lookup, scoped ownership, and graph-owned workgroup recomputation. Interleaved shader math and output are unchanged.

Key changes:
- `rotary_embedding_hf.wgsl` and generated registry entry — one thread per pair for HuggingFace rotate-half.
- `RotaryEmbedding.cpp` — named validation, typed resize contexts, shared dispatch descriptors, and one initial/resize grid picker per route.
- Native tests — malformed input rejection, dynamic bounds, shrink/regrow reuse, and HF/interleaved lifecycle coverage.

Co-authored-with: Claude Code.
ghstack-source-id: 411961455
@exported-using-ghexport

Differential Revision: [D113171746](https://our.internmc.facebook.com/intern/diff/D113171746/)
JCNTH added a commit that referenced this pull request Aug 7, 2026
Pull Request resolved: #21135

**Add HuggingFace rotate-half RoPE and migrate RoPE to shared dispatch construction**

Qwen models pair the first and second halves of each head vector rather than adjacent elements. This adds the rotate-half operator with dynamic start-position updates, full-dimension Q/K handling, generated WGSL, and strict frequency-table bounds.

The shared RoPE handler also closes the dispatch-boilerplate review: it uses `graph.device()`, typed graph-owned parameter buffers, descriptor-driven bindings and pipelines, named validation and resize callbacks, generated shader-registry lookup, scoped ownership, and graph-owned workgroup recomputation. Interleaved shader math and output are unchanged.

Key changes:
- `rotary_embedding_hf.wgsl` and generated registry entry — one thread per pair for HuggingFace rotate-half.
- `RotaryEmbedding.cpp` — named validation, typed resize contexts, shared dispatch descriptors, and one initial/resize grid picker per route.
- Native tests — malformed input rejection, dynamic bounds, shrink/regrow reuse, and HF/interleaved lifecycle coverage.

Co-authored-with: Claude Code.
ghstack-source-id: 411961455
@exported-using-ghexport

Differential Revision: [D113171746](https://our.internmc.facebook.com/intern/diff/D113171746/)
JCNTH added a commit that referenced this pull request Aug 7, 2026
Pull Request resolved: #21135

**Add HuggingFace rotate-half RoPE and migrate RoPE to shared dispatch construction**

Qwen models pair the first and second halves of each head vector rather than adjacent elements. This adds the rotate-half operator with dynamic start-position updates, full-dimension Q/K handling, generated WGSL, and strict frequency-table bounds.

The shared RoPE handler also closes the dispatch-boilerplate review: it uses `graph.device()`, typed graph-owned parameter buffers, descriptor-driven bindings and pipelines, named validation and resize callbacks, generated shader-registry lookup, scoped ownership, and graph-owned workgroup recomputation. Interleaved shader math and output are unchanged.

Key changes:
- `rotary_embedding_hf.wgsl` and generated registry entry — one thread per pair for HuggingFace rotate-half.
- `RotaryEmbedding.cpp` — named validation, typed resize contexts, shared dispatch descriptors, and one initial/resize grid picker per route.
- Native tests — malformed input rejection, dynamic bounds, shrink/regrow reuse, and HF/interleaved lifecycle coverage.

Co-authored-with: Claude Code.
ghstack-source-id: 411961455
@exported-using-ghexport

Differential Revision: [D113171746](https://our.internmc.facebook.com/intern/diff/D113171746/)
JCNTH added a commit that referenced this pull request Aug 7, 2026
Pull Request resolved: #21135

**Add HuggingFace rotate-half RoPE and migrate RoPE to shared dispatch construction**

Qwen models pair the first and second halves of each head vector rather than adjacent elements. This adds the rotate-half operator with dynamic start-position updates, full-dimension Q/K handling, generated WGSL, and strict frequency-table bounds.

The shared RoPE handler also closes the dispatch-boilerplate review: it uses `graph.device()`, typed graph-owned parameter buffers, descriptor-driven bindings and pipelines, named validation and resize callbacks, generated shader-registry lookup, scoped ownership, and graph-owned workgroup recomputation. Interleaved shader math and output are unchanged.

Key changes:
- `rotary_embedding_hf.wgsl` and generated registry entry — one thread per pair for HuggingFace rotate-half.
- `RotaryEmbedding.cpp` — named validation, typed resize contexts, shared dispatch descriptors, and one initial/resize grid picker per route.
- Native tests — malformed input rejection, dynamic bounds, shrink/regrow reuse, and HF/interleaved lifecycle coverage.

Co-authored-with: Claude Code.
ghstack-source-id: 411961455
@exported-using-ghexport

Differential Revision: [D113171746](https://our.internmc.facebook.com/intern/diff/D113171746/)
JCNTH added a commit that referenced this pull request Aug 7, 2026
Pull Request resolved: #21135

**Add HuggingFace rotate-half RoPE and migrate RoPE to shared dispatch construction**

Qwen models pair the first and second halves of each head vector rather than adjacent elements. This adds the rotate-half operator with dynamic start-position updates, full-dimension Q/K handling, generated WGSL, and strict frequency-table bounds.

The shared RoPE handler also closes the dispatch-boilerplate review: it uses `graph.device()`, typed graph-owned parameter buffers, descriptor-driven bindings and pipelines, named validation and resize callbacks, generated shader-registry lookup, scoped ownership, and graph-owned workgroup recomputation. Interleaved shader math and output are unchanged.

Key changes:
- `rotary_embedding_hf.wgsl` and generated registry entry — one thread per pair for HuggingFace rotate-half.
- `RotaryEmbedding.cpp` — named validation, typed resize contexts, shared dispatch descriptors, and one initial/resize grid picker per route.
- Native tests — malformed input rejection, dynamic bounds, shrink/regrow reuse, and HF/interleaved lifecycle coverage.

Co-authored-with: Claude Code.
ghstack-source-id: 411961455
@exported-using-ghexport

Differential Revision: [D113171746](https://our.internmc.facebook.com/intern/diff/D113171746/)
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. meta-exported

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants