Skip to content

feat(npu): add Qwen3.5-35B-A3B training scripts and NPU infrastructure - #63

Merged
NINGBENZHE merged 7 commits into
redai-studio:mainfrom
hbamboo:ascend-dev-0707
Jul 14, 2026
Merged

NINGBENZHE merged 7 commits into
redai-studio:mainfrom
hbamboo:ascend-dev-0707

Conversation

@hbamboo

@hbamboo hbamboo commented Jul 8, 2026 •

Copy link
Copy Markdown
Contributor

⭐ Feature

Qwen3.5-35B-A3B training scripts

  • Add 16xNPU fully-async training script with GRPO support
    • TP=4, EP=8, --num-iters-per-train-update 32, max-staleness=2
    • Actor (8 GPUs) + Rollout (8 GPUs) via independent placement groups
  • Add 16xNPU colocate training script
    • Actor and Rollout time-share 16 GPUs via --colocate
  • Add 8xNPU colocate training script (single-node)

Qwen3.5-9B script tuning

  • Colocate: add --max-actor-ckpt-to-keep 1 to limit disk usage
  • Async: tune SGLang params (mem-fraction 0.7→0.8, cuda-graph-bs refined, add --sglang-max-running-requests 256)

🔩 NPU Infrastructure

Docker

  • Pin TransferQueue to commit dcc78f0 for build reproducibility
  • Add protobuf==6.33.6 pin to avoid compatibility issues
  • Apply new sglang-npu.patch during SGLang build
  • Fix MindSpeed-Bridge clone indentation

NPU Patches

  • sglang-npu.patch (new): NPU-specific fixes for SGLang inference engine
  • megatron.patch: Add gate slicing fix for num_query_groups < world_size in attention module
  • mindspeed-bridge.patch: Refactor and expand NPU bridge compatibility

🐛 Bug Fix

  • Set partition_dim=0 for NPU in weight_update/common.py linear_fc1.weight all-gather path to work around Megatron grouped MoE bug

📝 Documentation

  • Update npu-training.md: add Qwen3.5-35B-A3B to model support table, expand device list to 16 cards, update roadmap

  • 300step results compare with H800:

  • colocate:

raw-reward-qwen35-35B - async: raw-reward-qwen35-35B-async

tang27-1 and others added 4 commits July 8, 2026 14:01
feat(npu): add 35B-A3B NPU training scripts

# ⭐ Feature

## Qwen3.5-35B-A3B training scripts

- Add 16xNPU async and colocate training scripts
- Add 8xNPU colocate training script
- Tune TP size, eval-interval, checkpoint paths, memory config
- Update NPU docker configuration and training documentation

## NPU infrastructure

- Add NPU docker patches (megatron, mindspeed-bridge, sglang-npu)
- Update weight_update/common.py for NPU compatibility
- Add and update NPU training documentation
feat(npu): add Qwen3.5-35B-A3B NPU training scripts

Created-by: Tgz27
Commit-by: Tgz27
Merged-by: dabuliu123
Description: ## What

<!-- What changes does this PR introduce? -->

## Why

<!-- Why are these changes needed? Link related issues with "Fixes redai-studio#123" or "Relates to #456". -->

## How

<!-- How do the changes work? Describe the technical approach. -->

## Testing

<!-- How were the changes tested? Include commands, test results, or screenshots. -->

- [ ] `pre-commit run --all-files` passes
- [ ] Tests pass (`pytest tests/`)
- [ ] New tests added (if applicable)
- [ ] Documentation updated (if applicable)

## Type of Change

- [ ] Bug fix (non-breaking change that fixes an issue)
- [ ] New feature (non-breaking change that adds functionality)
- [ ] Breaking change (fix or feature that would cause existing functionality to change)
- [ ] Documentation update
- [ ] Refactoring (no functional changes)
- [ ] Performance improvement
- [ ] CI/CD or build changes

## Screenshots / Logs

<!-- If applicable, add screenshots or log output to help explain the changes. -->


See merge request: hw-pbclouds/Relax!46
fix(scripts): correct 35B-A3B colocate script headers
fix(scripts): correct 35B-A3B colocate script headers

Created-by: Tgz27
Commit-by: Tgz27
Merged-by: hZhang111
Description: ## What

<!-- What changes does this PR introduce? -->

## Why

<!-- Why are these changes needed? Link related issues with "Fixes redai-studio#123" or "Relates to #456". -->

## How

<!-- How do the changes work? Describe the technical approach. -->

## Testing

<!-- How were the changes tested? Include commands, test results, or screenshots. -->

- [ ] `pre-commit run --all-files` passes
- [ ] Tests pass (`pytest tests/`)
- [ ] New tests added (if applicable)
- [ ] Documentation updated (if applicable)

## Type of Change

- [ ] Bug fix (non-breaking change that fixes an issue)
- [ ] New feature (non-breaking change that adds functionality)
- [ ] Breaking change (fix or feature that would cause existing functionality to change)
- [ ] Documentation update
- [ ] Refactoring (no functional changes)
- [ ] Performance improvement
- [ ] CI/CD or build changes

## Screenshots / Logs

<!-- If applicable, add screenshots or log output to help explain the changes. -->


See merge request: hw-pbclouds/Relax!47
if "linear_fc1.weight" in name and "vision_model" not in name:
param_partitions = [p.chunk(2, dim=0) for p in param_partitions]
param_partitions = [p[0] for p in param_partitions] + [p[1] for p in param_partitions]
if is_npu_available:

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

这块权重转换逻辑,npu为什么不一样?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

partition_dim gpu是0,npu是1,导致shape不匹配,先修改规避

# bash scripts/training/text/run-qwen35-35B-A3B-16xgpu-async.sh
set -ex
set -o pipefail
export HCCL_SOCKET_IFNAME=enp23s0f3

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

这种网卡的参数配置不要写死,改成类似
export HCCL_SOCKET_IFNAME="${HCCL_SOCKET_IFNAME:-enp23s0f3}"


ulimit -n 65535

export HCCL_SOCKET_IFNAME=enp23s0f3

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

同上

tang27-1 and others added 3 commits July 9, 2026 16:38
# 🐛 Bug Fix

## MindSpeed bridge patch

- Remove finalize() and validate_parallelism() from qwen35_vl_provider.py
  instead of commenting them out — cleaner patch with less unused code
- Clean up sglang-npu.patch

## NPU weight update

- Add TODO comment for partition_dim=0 NPU workaround in common.py

## Training scripts

- Make HCCL_SOCKET_IFNAME, GLOO_SOCKET_IFNAME, TP_SOCKET_IFNAME
  overridable via environment variables (default: enp23s0f3) in both
  16xnpu async and colocate scripts
fix(npu): remove validate_parallelism, clean patches, tune 35B scripts

Created-by: Tgz27
Commit-by: Tgz27
Merged-by: dabuliu123
Description: Description:


## Summary

Fix NPU-related issues in MindSpeed bridge patch and 35B training scripts.

## Changes

- Remove finalize() and validate_parallelism() from qwen35_vl_provider.py instead of commenting out
- Clean up sglang-npu.patch
- Add TODO comment for partition_dim=0 NPU workaround in common.py
- Make HCCL/GLOO/TP socket ifname overridable via env vars in 35B scripts

See merge request: hw-pbclouds/Relax!48

@NINGBENZHE NINGBENZHE left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

/lgtm

@NINGBENZHE
NINGBENZHE merged commit 727e0d3 into redai-studio:main Jul 14, 2026
5 checks passed
ZiyiTsang pushed a commit to ZiyiTsang/Relax that referenced this pull request Aug 6, 2026
redai-studio#63)

# ⭐ Feature
## Qwen3.5-35B-A3B training scripts
- Add 16xNPU fully-async training script with GRPO support
  - TP=4, EP=8, `--num-iters-per-train-update 32`, max-staleness=2
  - Actor (8 GPUs) + Rollout (8 GPUs) via independent placement groups
- Add 16xNPU colocate training script
  - Actor and Rollout time-share 16 GPUs via `--colocate`
- Add 8xNPU colocate training script (single-node)
## Qwen3.5-9B script tuning
- Colocate: add `--max-actor-ckpt-to-keep 1` to limit disk usage
- Async: tune SGLang params (mem-fraction 0.7→0.8, cuda-graph-bs
refined, add `--sglang-max-running-requests 256`)
---
# 🔩 NPU Infrastructure
## Docker
- Pin TransferQueue to commit `dcc78f0` for build reproducibility
- Add `protobuf==6.33.6` pin to avoid compatibility issues
- Apply new `sglang-npu.patch` during SGLang build
- Fix MindSpeed-Bridge clone indentation
## NPU Patches
- **sglang-npu.patch** (new): NPU-specific fixes for SGLang inference
engine
- **megatron.patch**: Add gate slicing fix for `num_query_groups <
world_size` in attention module
- **mindspeed-bridge.patch**: Refactor and expand NPU bridge
compatibility
---
# 🐛 Bug Fix
- Set `partition_dim=0` for NPU in `weight_update/common.py`
`linear_fc1.weight` all-gather path to work around Megatron grouped MoE
bug
---
# 📝 Documentation
- Update `npu-training.md`: add Qwen3.5-35B-A3B to model support table,
expand device list to 16 cards, update roadmap

- 300step results compare with H800:
 - colocate:
<img width="984" height="1146" alt="raw-reward-qwen35-35B"
src="https://github.com/user-attachments/assets/69ebcf0e-7dda-43d9-9dd6-0def06677629"
/>
- async:

<img width="984" height="1096" alt="raw-reward-qwen35-35B-async"
src="https://github.com/user-attachments/assets/381995bb-863d-482f-be3f-27196d2ac66a"
/>

---------

Co-authored-by: Tgz27 <617796318@qq.com>
Co-authored-by: dabuliu123 <270334047@qq.com>
Co-authored-by: hZhang111 <hZhang111@noreply.gitcode.com>
Co-authored-by: 宁本哲 <ningbenzhe@xiaohongshu.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants