Repository navigation
ci(evals): add AMD legs to the BFCL A/B, gated on the mean of four runs - #2832
Draft
chunfangamd wants to merge 1 commit into
Draft
chunfangamd wants to merge 1 commit into
chunfangamd wants to merge 1 commit into
Conversation
Add gpt-oss-amd and qwen3.8-amd to nightly-bfcl.yml: the H100 legs' models and parsers on an AMD Instinct self-hosted runner, with vLLM from the ROCm wheel of the CI pin (scripts/ci_install_vllm_rocm.sh). Each arm fits one GPU, so a leg runs four A/B pairs side by side on one node (arm_mode: repeated, scripts/bfcl/run_repeats.sh) and gates on their mean (run_ab.py --combine). The legs stay off until the runner is registered and SMG_RUN_AMD_LEGS is "true", and never run on pull_request. Also add a ROCm GPU cleanup to ci_killall_sglang.sh (rocm nuke_gpus), and set BFCL_RUN_DIR for the always() teardown step, which looked for pidfiles in the wrong directory. Signed-off-by: Chun Fang <chun.fang@amd.com>
|
Important Draft PR not reviewedDraft PRs are not automatically reviewed by default.
To automatically review draft PRs, update your CodeRabbit configuration: reviews:
auto_review:
drafts: true
Comment |
3 tasks
| inputs: | ||
| only: | ||
| description: "Run only this matrix leg (qwen3.8|gpt-oss|deepseek-v4.1|minimax-m3|glm-5.3-flash); empty = all" | ||
| description: "Run only this matrix leg (qwen3.8|gpt-oss|deepseek-v4.1|minimax-m3|glm-5.3-flash|gpt-oss-amd|qwen3.8-amd); empty = all" |
There was a problem hiding this comment.
should this be in here or add another variable to define hardware of nvidia or amd or both with default being both?
Comment on lines
+126
to
+128
| # AMD legs: the AMD runner's label, and "true" once that runner is online. | ||
| AMD_RUNNER: ${{ vars.SMG_RUNNER_AMD_GPU_8 }} | ||
| AMD_ENABLED: ${{ vars.SMG_RUN_AMD_LEGS }} |
There was a problem hiding this comment.
has an stable fleet of AMD CI runners been hooked up into the upstream smg repo yet?
| @@ -0,0 +1,101 @@ | |||
| #!/usr/bin/env bash | |||
There was a problem hiding this comment.
what the need for this verus on existing NVIDIA? can we keep scope to just amd enablement
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Part of #2736, which asks for the BFCL nightly to cover AMD GPUs. This PR adds two AMD legs to
nightly-bfcl.ymlthat run the H100 legs' two models, gpt-oss-120b and Qwen3.8-27B, and keeps them off until AMD's self-hosted runner is registered.What the AMD legs do
wheels.vllm.ai/rocm(newscripts/ci_install_vllm_rocm.sh), so an AMD leg and its H100 counterpart differ only in hardware. The SMG wheel,smg-grpc-protoandsmg-grpc-servicercome from the same build and source as for the H100 legs.arm_mode: repeated,scripts/bfcl/run_repeats.shandrun_ab.py --combine. The job summary shows the mean table, then each run's overall result and the range.Why four runs
BFCL samples: vLLM raises BFCL's temperature of 0.001 to 0.01, and the model's default
top_p/top_kapply. We ran Qwen3-0.6B twice with identical settings, once onsimple_python,irrelevanceand once on the 7 non-live categories. Between the two runs, the same arm changed its verdict on 2–13% of the cases per category, and the unweighted Δ moved from +0.83 to −0.42 on the first set and from −1.01 to +0.44 on the second. That leaves little margin under a 2-point tolerance. In four identical runs with this PR's setup (see Validation), the unweighted Δ ranged from −0.74 to +2.33. Greedy decoding does not remove the noise either: at temperature 0 the same arm still changed its verdict on about as many cases between runs, so it comes from batching concurrent requests, not from sampling. Averaging four runs halves the noise, and because each run needs only two of the node's eight GPUs, it adds no wall-clock time.Switching the AMD legs on
Until then, the AMD legs run only from a
workflow_dispatchthat names one withonly, and never onpull_request, so the runner only ever takes code frommain. To switch them on:8-gpu-mi325x, or set the repository variableSMG_RUNNER_AMD_GPU_8to its label.SMG_RUN_AMD_LEGStotrue.The runner needs ROCm 7.2 installed, glibc 2.34 or newer, network access to PyPI,
wheels.vllm.aiand Hugging Face, the two models under/models(or room to download them), andlsoffor the GPU cleanup. The ROCm release matters because the wheels are built for 7.2 and load some of its system libraries;scripts/ci_install_vllm_rocm.shchecks it before installing.scripts/ci_setup_python_venv.shprovisions Python 3.12, which the ROCm wheels require.Other changes
ci_killall_sglang.sh rocm nuke_gpusnow kills the processes holding/dev/kfd, the ROCm counterpart ofnuke_gpus. The CUDA path is unchanged.always()teardown step now setsBFCL_RUN_DIR. Without it,launch_arm.sh stoplooked for pidfiles in/tmp/bfcl_abinstead of$RUNNER_TEMP/bfcl_run, so on every leg this backstop never found the servers it is meant to stop.scripts/bfcl/README.mdandCONTRIBUTING.mddescribe the AMD legs and the two variables.Validation
run_ab.py --combinehas four new unit tests; all 7 tests intest_bfcl_run_ab.pypass, and ruff and pre-commit pass.We ran the matrix script for scheduled, dispatched and pull-request runs, with the AMD variables set and unset. The H100 and Blackwell legs come out exactly as before; the AMD legs appear only as described above.
On an MI355X node, with the same scripts an AMD runner would run (
ci_setup_python_venv.sh,ci_install_vllm_rocm.sh, the SMG wheel andbfcl-eval, thenrun_repeats.sh). The first attempt, on a host whose system ROCm is 7.1.1, failed when torch importedamdsmi, which needs 7.2'slibamd_smi; that is why the install script now checks the ROCm release first. The second ran with ROCm 7.2.3 libraries; the setup took about 2 minutes. Scores are vLLM / SMG, with Δ = SMG − vLLM:gpt-oss-amdas configured, four runs of the 17 categories: the eight servers were up in about 8 minutes (loading from network storage) and scoring took 41 minutes. The mean of the four runs was 52.70 / 52.23 (−0.47) unweighted and 66.75 / 66.40 (−0.34) weighted, and the runs' unweighted Δ were −0.03, −0.11, −1.25 and −0.48. The 2026-10-05 H100 run, with the same vLLM 0.27.1, scored 52.73 / 52.00 and 67.10 / 66.22.qwen3.8-amd's setup onsimple_python,irrelevance, four runs: the mean was 86.44 / 86.12 (−0.31) unweighted and 87.97 / 87.73 (−0.23) weighted, with run Δ of −0.13, +0.04, −1.08 and −0.08 unweighted. On these two categories the H100 run scored 86.38 / 86.42 and 87.97 / 87.81. The slowest single case took about 20 minutes, which is why the leg keeps the H100 leg's 5-hour cap.The MI325X is gfx942 and the MI355X gfx950, so the first dispatch on the MI325X runner is the check for that hardware.
For reviewers
run_repeats.shworks for any leg whose arms leave room on its node; the H100 legs, at TP=2 on a 4-GPU runner, do not.