feat(npu): add Qwen3.5-35B-A3B training scripts and NPU infrastructure - #63
Merged
Merged
Conversation
feat(npu): add 35B-A3B NPU training scripts # ⭐ Feature ## Qwen3.5-35B-A3B training scripts - Add 16xNPU async and colocate training scripts - Add 8xNPU colocate training script - Tune TP size, eval-interval, checkpoint paths, memory config - Update NPU docker configuration and training documentation ## NPU infrastructure - Add NPU docker patches (megatron, mindspeed-bridge, sglang-npu) - Update weight_update/common.py for NPU compatibility - Add and update NPU training documentation
feat(npu): add Qwen3.5-35B-A3B NPU training scripts Created-by: Tgz27 Commit-by: Tgz27 Merged-by: dabuliu123 Description: ## What <!-- What changes does this PR introduce? --> ## Why <!-- Why are these changes needed? Link related issues with "Fixes redai-studio#123" or "Relates to #456". --> ## How <!-- How do the changes work? Describe the technical approach. --> ## Testing <!-- How were the changes tested? Include commands, test results, or screenshots. --> - [ ] `pre-commit run --all-files` passes - [ ] Tests pass (`pytest tests/`) - [ ] New tests added (if applicable) - [ ] Documentation updated (if applicable) ## Type of Change - [ ] Bug fix (non-breaking change that fixes an issue) - [ ] New feature (non-breaking change that adds functionality) - [ ] Breaking change (fix or feature that would cause existing functionality to change) - [ ] Documentation update - [ ] Refactoring (no functional changes) - [ ] Performance improvement - [ ] CI/CD or build changes ## Screenshots / Logs <!-- If applicable, add screenshots or log output to help explain the changes. --> See merge request: hw-pbclouds/Relax!46
fix(scripts): correct 35B-A3B colocate script headers
fix(scripts): correct 35B-A3B colocate script headers Created-by: Tgz27 Commit-by: Tgz27 Merged-by: hZhang111 Description: ## What <!-- What changes does this PR introduce? --> ## Why <!-- Why are these changes needed? Link related issues with "Fixes redai-studio#123" or "Relates to #456". --> ## How <!-- How do the changes work? Describe the technical approach. --> ## Testing <!-- How were the changes tested? Include commands, test results, or screenshots. --> - [ ] `pre-commit run --all-files` passes - [ ] Tests pass (`pytest tests/`) - [ ] New tests added (if applicable) - [ ] Documentation updated (if applicable) ## Type of Change - [ ] Bug fix (non-breaking change that fixes an issue) - [ ] New feature (non-breaking change that adds functionality) - [ ] Breaking change (fix or feature that would cause existing functionality to change) - [ ] Documentation update - [ ] Refactoring (no functional changes) - [ ] Performance improvement - [ ] CI/CD or build changes ## Screenshots / Logs <!-- If applicable, add screenshots or log output to help explain the changes. --> See merge request: hw-pbclouds/Relax!47
NINGBENZHE
reviewed
Jul 9, 2026
| if "linear_fc1.weight" in name and "vision_model" not in name: | ||
| param_partitions = [p.chunk(2, dim=0) for p in param_partitions] | ||
| param_partitions = [p[0] for p in param_partitions] + [p[1] for p in param_partitions] | ||
| if is_npu_available: |
Contributor
Author
There was a problem hiding this comment.
partition_dim gpu是0,npu是1,导致shape不匹配,先修改规避
| # bash scripts/training/text/run-qwen35-35B-A3B-16xgpu-async.sh | ||
| set -ex | ||
| set -o pipefail | ||
| export HCCL_SOCKET_IFNAME=enp23s0f3 |
Member
There was a problem hiding this comment.
这种网卡的参数配置不要写死,改成类似
export HCCL_SOCKET_IFNAME="${HCCL_SOCKET_IFNAME:-enp23s0f3}"
|
|
||
| ulimit -n 65535 | ||
|
|
||
| export HCCL_SOCKET_IFNAME=enp23s0f3 |
# 🐛 Bug Fix ## MindSpeed bridge patch - Remove finalize() and validate_parallelism() from qwen35_vl_provider.py instead of commenting them out — cleaner patch with less unused code - Clean up sglang-npu.patch ## NPU weight update - Add TODO comment for partition_dim=0 NPU workaround in common.py ## Training scripts - Make HCCL_SOCKET_IFNAME, GLOO_SOCKET_IFNAME, TP_SOCKET_IFNAME overridable via environment variables (default: enp23s0f3) in both 16xnpu async and colocate scripts
fix(npu): remove validate_parallelism, clean patches, tune 35B scripts Created-by: Tgz27 Commit-by: Tgz27 Merged-by: dabuliu123 Description: Description: ## Summary Fix NPU-related issues in MindSpeed bridge patch and 35B training scripts. ## Changes - Remove finalize() and validate_parallelism() from qwen35_vl_provider.py instead of commenting out - Clean up sglang-npu.patch - Add TODO comment for partition_dim=0 NPU workaround in common.py - Make HCCL/GLOO/TP socket ifname overridable via env vars in 35B scripts See merge request: hw-pbclouds/Relax!48
ZiyiTsang
pushed a commit
to ZiyiTsang/Relax
that referenced
this pull request
Aug 6, 2026
redai-studio#63) # ⭐ Feature ## Qwen3.5-35B-A3B training scripts - Add 16xNPU fully-async training script with GRPO support - TP=4, EP=8, `--num-iters-per-train-update 32`, max-staleness=2 - Actor (8 GPUs) + Rollout (8 GPUs) via independent placement groups - Add 16xNPU colocate training script - Actor and Rollout time-share 16 GPUs via `--colocate` - Add 8xNPU colocate training script (single-node) ## Qwen3.5-9B script tuning - Colocate: add `--max-actor-ckpt-to-keep 1` to limit disk usage - Async: tune SGLang params (mem-fraction 0.7→0.8, cuda-graph-bs refined, add `--sglang-max-running-requests 256`) --- # 🔩 NPU Infrastructure ## Docker - Pin TransferQueue to commit `dcc78f0` for build reproducibility - Add `protobuf==6.33.6` pin to avoid compatibility issues - Apply new `sglang-npu.patch` during SGLang build - Fix MindSpeed-Bridge clone indentation ## NPU Patches - **sglang-npu.patch** (new): NPU-specific fixes for SGLang inference engine - **megatron.patch**: Add gate slicing fix for `num_query_groups < world_size` in attention module - **mindspeed-bridge.patch**: Refactor and expand NPU bridge compatibility --- # 🐛 Bug Fix - Set `partition_dim=0` for NPU in `weight_update/common.py` `linear_fc1.weight` all-gather path to work around Megatron grouped MoE bug --- # 📝 Documentation - Update `npu-training.md`: add Qwen3.5-35B-A3B to model support table, expand device list to 16 cards, update roadmap - 300step results compare with H800: - colocate: <img width="984" height="1146" alt="raw-reward-qwen35-35B" src="https://github.com/user-attachments/assets/69ebcf0e-7dda-43d9-9dd6-0def06677629" /> - async: <img width="984" height="1096" alt="raw-reward-qwen35-35B-async" src="https://github.com/user-attachments/assets/381995bb-863d-482f-be3f-27196d2ac66a" /> --------- Co-authored-by: Tgz27 <617796318@qq.com> Co-authored-by: dabuliu123 <270334047@qq.com> Co-authored-by: hZhang111 <hZhang111@noreply.gitcode.com> Co-authored-by: 宁本哲 <ningbenzhe@xiaohongshu.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
⭐ Feature
Qwen3.5-35B-A3B training scripts
--num-iters-per-train-update 32, max-staleness=2--colocateQwen3.5-9B script tuning
--max-actor-ckpt-to-keep 1to limit disk usage--sglang-max-running-requests 256)🔩 NPU Infrastructure
Docker
dcc78f0for build reproducibilityprotobuf==6.33.6pin to avoid compatibility issuessglang-npu.patchduring SGLang buildNPU Patches
num_query_groups < world_sizein attention module🐛 Bug Fix
partition_dim=0for NPU inweight_update/common.pylinear_fc1.weightall-gather path to work around Megatron grouped MoE bug📝 Documentation
Update
npu-training.md: add Qwen3.5-35B-A3B to model support table, expand device list to 16 cards, update roadmap300step results compare with H800:
colocate: