Add parallel chunked prefill for channelwise gated delta rule (#21061)#21061
Conversation
🔗 Helpful Links🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/21061
Note: Links to docs will display an error until the docs builds have been completed. ✅ No FailuresAs of commit 1b7ba84 with merge base 179c4ee ( This comment was automatically generated by Dr. CI and updates every 15 minutes. |
|
@JakeStevens has exported this pull request. If you are a Meta employee, you can view the originating Diff in D112597348. |
This PR needs a
|
…h#21061) Summary: Route the channelwise gated delta rule by sequence length: T == 1 keeps the two-pass token recurrence for autoregressive decode, while T != 1 uses a chunkwise WY/UT formulation for prefill. The chunked path computes per-channel log-decay prefixes, causal query/key terms, the beta-folded triangular transform, WY pseudo-keys and pseudo-values, and inter-chunk state carry. It handles a ragged final chunk without a separate tail implementation. Parallelize independent (batch, head) work across the ExecuTorch threadpool. Each worker receives a disjoint slice of one temporary scratch arena, avoiding shared mutable buffers while amortizing allocation across chunks. Reviewed By: billmguo Differential Revision: D112597348
1509cce to
85d15b3
Compare
…h#21061) Summary: Route the channelwise gated delta rule by sequence length: T == 1 keeps the two-pass token recurrence for autoregressive decode, while T != 1 uses a chunkwise WY/UT formulation for prefill. The chunked path computes per-channel log-decay prefixes, causal query/key terms, the beta-folded triangular transform, WY pseudo-keys and pseudo-values, and inter-chunk state carry. It handles a ragged final chunk without a separate tail implementation. Parallelize independent (batch, head) work across the ExecuTorch threadpool. Each worker receives a disjoint slice of one temporary scratch arena, avoiding shared mutable buffers while amortizing allocation across chunks. Reviewed By: billmguo Differential Revision: D112597348
…BUCK (pytorch#21105) Summary: Relands the two-pass optimization for the `channelwise_gated_delta_rule` custom op (originally pytorch#21020, D112596724), which was reverted in D113048961 because it broke OSS `unittest macos / linux`. The revert was caused by the benchmark BUCK target: ``` runtime.python_binary(name = ..., srcs = [...], main_module = ...) ``` Fix: move the source into a `runtime.python_library` and have the `runtime.python_binary` reference it via `deps` with only `main_module` Differential Revision: D113076546
…h#21061) Summary: Route the channelwise gated delta rule by sequence length: T == 1 keeps the two-pass token recurrence for autoregressive decode, while T != 1 uses a chunkwise WY/UT formulation for prefill. The chunked path computes per-channel log-decay prefixes, causal query/key terms, the beta-folded triangular transform, WY pseudo-keys and pseudo-values, and inter-chunk state carry. It handles a ragged final chunk without a separate tail implementation. Parallelize independent (batch, head) work across the ExecuTorch threadpool. Each worker receives a disjoint slice of one temporary scratch arena, avoiding shared mutable buffers while amortizing allocation across chunks. Reviewed By: billmguo Differential Revision: D112597348
85d15b3 to
abb6ab8
Compare
…h#21061) Summary: Route the channelwise gated delta rule by sequence length: T == 1 keeps the two-pass token recurrence for autoregressive decode, while T != 1 uses a chunkwise WY/UT formulation for prefill. The chunked path computes per-channel log-decay prefixes, causal query/key terms, the beta-folded triangular transform, WY pseudo-keys and pseudo-values, and inter-chunk state carry. It handles a ragged final chunk without a separate tail implementation. Parallelize independent (batch, head) work across the ExecuTorch threadpool. Each worker receives a disjoint slice of one temporary scratch arena, avoiding shared mutable buffers while amortizing allocation across chunks. Reviewed By: billmguo Differential Revision: D112597348
Summary:
Route the channelwise gated delta rule by sequence length: T == 1 keeps the two-pass token recurrence for autoregressive decode, while T != 1 uses a chunkwise WY/UT formulation for prefill.
The chunked path computes per-channel log-decay prefixes, causal query/key terms, the beta-folded triangular transform, WY pseudo-keys and pseudo-values, and inter-chunk state carry. It handles a ragged final chunk without a separate tail implementation.
Parallelize independent (batch, head) work across the ExecuTorch threadpool. Each worker receives a disjoint slice of one temporary scratch arena, avoiding shared mutable buffers while amortizing allocation across chunks.
Reviewed By: billmguo
Differential Revision: D112597348