You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Commit 101e6af
Browse filesBrowse the repository at this point in the historyBrowse files
issue/1565 fix(nvidia): complete runtime support for Qwen MTP
Map E4M3 and BOOL through the existing ATen adaptor and preserve the
caller's CUDA device across NCCL communicator destruction.
Reuse the existing paged Prefill warp kernel for NVIDIA head size 256,
without changing other vendors' default dispatch. Extend existing
multi-page/long-context coverage and add finite FP8/mask cast checks.
Validation: fresh SM86 build, 88 paged Prefill cases, 2 cast tests,
and TP2 communicator teardown from both caller devices.
Closes#1565
0 commit comments