[ET-VK][sdpa] Vendor-adaptive head_dim output-tiling (TILE_N4=2) in GQA AV coop-GEMV#21064
Conversation
🔗 Helpful Links🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/21064
Note: Links to docs will display an error until the docs builds have been completed. ❗ 1 Active SEVsThere are 1 currently active SEVs. If your PR is affected, please view them below: ⏳ No Failures, 2 PendingAs of commit 8bea766 with merge base 37400d9 ( This comment was automatically generated by Dr. CI and updates every 15 minutes. |
This PR needs a
|
digantdesai
left a comment
There was a problem hiding this comment.
Review automatically exported from Phabricator review in Meta.
…QA AV coop-GEMV Pull Request resolved: #21064 Builds on the GQA-reuse AV coop-GEMV (parent commit). The base `sdpa_compute_out_gqa_coop` shader assigns one head_dim texel (TILE_N4=1) per workgroup x-slot. This adds a head_dim output-tiled variant, `sdpa_compute_out_gqa_coop_tile2` (TILE_N4=2): each workgroup owns 2 head_dim texels and emits G x TILE_N4 outputs, so the per-context-texel attn-weight loads and the shared-memory tree reduction are amortized over twice as many outputs. The AV coop GEMV GLSL is already a TILE_N4-parameterized template (the `partial_n_tile` uniform branch handles the D4 % TILE_N4 != 0 tail), so this change is purely a new codegen variant plus dispatch wiring — no shader-body change. Selection is vendor-adaptive: `pick_sdpa_av_shader` appends the `_tile2` suffix only on Adreno (`graph->device_is_adreno()`). Tiling is a consistent win on Adreno (AV ~1.14-1.63x over the base GQA variant) but a regression on Mali at common decode contexts (~0.67-0.86x, interleaved median-of-N), so Mali and other vendors keep the base `sdpa_compute_out_gqa_coop`. `pick_sdpa_av_global_wg_size` keys off the same `_tile2` suffix: the tiled variant's x-dim collapses from D4 to div_up(D4, 2) since each workgroup now covers 2 head_dim texels. Because the variant is Adreno-only in production, the test-only `gqa_override` knob (see SDPA.h) is extended so tests can pin the variant on any device: `kGqaOverrideForceTile2` / `kGqaOverrideForceBase` force the tiled / base variant regardless of vendor (via `resolve_use_tile2`), giving the tiled shader deterministic coverage on Mali / SwiftShader. `resolve_use_gqa` is unchanged (any non-`ForceNonGqa` value still forces the GQA family, VK_CHECK'd for eligibility). ghstack-source-id: 405400514 @exported-using-ghexport Differential Revision: [D112906313](https://our.internmc.facebook.com/intern/diff/D112906313/)
5ee2d71
into
gh/SS-JIA/577/base
…QA AV coop-GEMV Pull Request resolved: #21064 Builds on the GQA-reuse AV coop-GEMV (parent commit). The base `sdpa_compute_out_gqa_coop` shader assigns one head_dim texel (TILE_N4=1) per workgroup x-slot. This adds a head_dim output-tiled variant, `sdpa_compute_out_gqa_coop_tile2` (TILE_N4=2): each workgroup owns 2 head_dim texels and emits G x TILE_N4 outputs, so the per-context-texel attn-weight loads and the shared-memory tree reduction are amortized over twice as many outputs. The AV coop GEMV GLSL is already a TILE_N4-parameterized template (the `partial_n_tile` uniform branch handles the D4 % TILE_N4 != 0 tail), so this change is purely a new codegen variant plus dispatch wiring — no shader-body change. Selection is vendor-adaptive: `pick_sdpa_av_shader` appends the `_tile2` suffix only on Adreno (`graph->device_is_adreno()`). Tiling is a consistent win on Adreno (AV ~1.14-1.63x over the base GQA variant) but a regression on Mali at common decode contexts (~0.67-0.86x, interleaved median-of-N), so Mali and other vendors keep the base `sdpa_compute_out_gqa_coop`. `pick_sdpa_av_global_wg_size` keys off the same `_tile2` suffix: the tiled variant's x-dim collapses from D4 to div_up(D4, 2) since each workgroup now covers 2 head_dim texels. Because the variant is Adreno-only in production, the test-only `gqa_override` knob (see SDPA.h) is extended so tests can pin the variant on any device: `kGqaOverrideForceTile2` / `kGqaOverrideForceBase` force the tiled / base variant regardless of vendor (via `resolve_use_tile2`), giving the tiled shader deterministic coverage on Mali / SwiftShader. `resolve_use_gqa` is unchanged (any non-`ForceNonGqa` value still forces the GQA family, VK_CHECK'd for eligibility). ghstack-source-id: 405400514 @exported-using-ghexport Differential Revision: [D112906313](https://our.internmc.facebook.com/intern/diff/D112906313/)
…QA AV coop-GEMV Pull Request resolved: #21064 Builds on the GQA-reuse AV coop-GEMV (parent commit). The base `sdpa_compute_out_gqa_coop` shader assigns one head_dim texel (TILE_N4=1) per workgroup x-slot. This adds a head_dim output-tiled variant, `sdpa_compute_out_gqa_coop_tile2` (TILE_N4=2): each workgroup owns 2 head_dim texels and emits G x TILE_N4 outputs, so the per-context-texel attn-weight loads and the shared-memory tree reduction are amortized over twice as many outputs. The AV coop GEMV GLSL is already a TILE_N4-parameterized template (the `partial_n_tile` uniform branch handles the D4 % TILE_N4 != 0 tail), so this change is purely a new codegen variant plus dispatch wiring — no shader-body change. Selection is vendor-adaptive: `pick_sdpa_av_shader` appends the `_tile2` suffix only on Adreno (`graph->device_is_adreno()`). Tiling is a consistent win on Adreno (AV ~1.14-1.63x over the base GQA variant) but a regression on Mali at common decode contexts (~0.67-0.86x, interleaved median-of-N), so Mali and other vendors keep the base `sdpa_compute_out_gqa_coop`. `pick_sdpa_av_global_wg_size` keys off the same `_tile2` suffix: the tiled variant's x-dim collapses from D4 to div_up(D4, 2) since each workgroup now covers 2 head_dim texels. Because the variant is Adreno-only in production, the test-only `gqa_override` knob (see SDPA.h) is extended so tests can pin the variant on any device: `kGqaOverrideForceTile2` / `kGqaOverrideForceBase` force the tiled / base variant regardless of vendor (via `resolve_use_tile2`), giving the tiled shader deterministic coverage on Mali / SwiftShader. `resolve_use_gqa` is unchanged (any non-`ForceNonGqa` value still forces the GQA family, VK_CHECK'd for eligibility). ghstack-source-id: 405400514 @exported-using-ghexport Differential Revision: [D112906313](https://our.internmc.facebook.com/intern/diff/D112906313/)
…QA AV coop-GEMV Pull Request resolved: #21064 Builds on the GQA-reuse AV coop-GEMV (parent commit). The base `sdpa_compute_out_gqa_coop` shader assigns one head_dim texel (TILE_N4=1) per workgroup x-slot. This adds a head_dim output-tiled variant, `sdpa_compute_out_gqa_coop_tile2` (TILE_N4=2): each workgroup owns 2 head_dim texels and emits G x TILE_N4 outputs, so the per-context-texel attn-weight loads and the shared-memory tree reduction are amortized over twice as many outputs. The AV coop GEMV GLSL is already a TILE_N4-parameterized template (the `partial_n_tile` uniform branch handles the D4 % TILE_N4 != 0 tail), so this change is purely a new codegen variant plus dispatch wiring — no shader-body change. Selection is vendor-adaptive: `pick_sdpa_av_shader` appends the `_tile2` suffix only on Adreno (`graph->device_is_adreno()`). Tiling is a consistent win on Adreno (AV ~1.14-1.63x over the base GQA variant) but a regression on Mali at common decode contexts (~0.67-0.86x, interleaved median-of-N), so Mali and other vendors keep the base `sdpa_compute_out_gqa_coop`. `pick_sdpa_av_global_wg_size` keys off the same `_tile2` suffix: the tiled variant's x-dim collapses from D4 to div_up(D4, 2) since each workgroup now covers 2 head_dim texels. Because the variant is Adreno-only in production, the test-only `gqa_override` knob (see SDPA.h) is extended so tests can pin the variant on any device: `kGqaOverrideForceTile2` / `kGqaOverrideForceBase` force the tiled / base variant regardless of vendor (via `resolve_use_tile2`), giving the tiled shader deterministic coverage on Mali / SwiftShader. `resolve_use_gqa` is unchanged (any non-`ForceNonGqa` value still forces the GQA family, VK_CHECK'd for eligibility). ghstack-source-id: 405400514 @exported-using-ghexport Differential Revision: [D112906313](https://our.internmc.facebook.com/intern/diff/D112906313/)
Stack from ghstack (oldest at bottom):
Builds on the GQA-reuse AV coop-GEMV (parent commit). The base
sdpa_compute_out_gqa_coopshader assigns one head_dim texel (TILE_N4=1) perworkgroup x-slot. This adds a head_dim output-tiled variant,
sdpa_compute_out_gqa_coop_tile2(TILE_N4=2): each workgroup owns 2 head_dimtexels and emits G x TILE_N4 outputs, so the per-context-texel attn-weight loads
and the shared-memory tree reduction are amortized over twice as many outputs.
The AV coop GEMV GLSL is already a TILE_N4-parameterized template (the
partial_n_tileuniform branch handles the D4 % TILE_N4 != 0 tail), so thischange is purely a new codegen variant plus dispatch wiring — no shader-body
change.
Selection is vendor-adaptive:
pick_sdpa_av_shaderappends the_tile2suffixonly on Adreno (
graph->device_is_adreno()). Tiling is a consistent win onAdreno (AV ~1.14-1.63x over the base GQA variant) but a regression on Mali at
common decode contexts (~0.67-0.86x, interleaved median-of-N), so Mali and other
vendors keep the base
sdpa_compute_out_gqa_coop.pick_sdpa_av_global_wg_sizekeys off the same_tile2suffix: the tiledvariant's x-dim collapses from D4 to div_up(D4, 2) since each workgroup now
covers 2 head_dim texels.
Because the variant is Adreno-only in production, the test-only
gqa_overrideknob (see SDPA.h) is extended so tests can pin the variant on any device:
kGqaOverrideForceTile2/kGqaOverrideForceBaseforce the tiled / basevariant regardless of vendor (via
resolve_use_tile2), giving the tiled shaderdeterministic coverage on Mali / SwiftShader.
resolve_use_gqais unchanged(any non-
ForceNonGqavalue still forces the GQA family, VK_CHECK'd foreligibility).
Differential Revision: D112906313