Skip to content

ci(evals): add AMD legs to the BFCL A/B, gated on the mean of four runs - #2832

Draft
chunfangamd wants to merge 1 commit into
smg-project:mainfrom
chunfangamd:chunfangamd/bfcl-amd-legs
Draft

chunfangamd wants to merge 1 commit into
smg-project:mainfrom
chunfangamd:chunfangamd/bfcl-amd-legs

Conversation

@chunfangamd

@chunfangamd chunfangamd commented Oct 6, 2026 •

Copy link
Copy Markdown
Contributor

Part of #2736, which asks for the BFCL nightly to cover AMD GPUs. This PR adds two AMD legs to nightly-bfcl.yml that run the H100 legs' two models, gpt-oss-120b and Qwen3.8-27B, and keeps them off until AMD's self-hosted runner is registered.

What the AMD legs do

  • They run in the same workflow run as the H100 legs, so they test the same commit on the same schedule (Mondays 07:17 UTC), with the same categories, sampling, tolerance, job summary and artifacts. Their failures reach the existing nightly-triage issue.
  • vLLM is the ROCm wheel of the CI pin, 0.27.1 from wheels.vllm.ai/rocm (new scripts/ci_install_vllm_rocm.sh), so an AMD leg and its H100 counterpart differ only in hardware. The SMG wheel, smg-grpc-proto and smg-grpc-servicer come from the same build and source as for the H100 legs.
  • Each arm fits on one GPU (an MI325X has 256 GB), so a leg runs four independent A/B pairs side by side on one 8-GPU node and gates on their mean: new arm_mode: repeated, scripts/bfcl/run_repeats.sh and run_ab.py --combine. The job summary shows the mean table, then each run's overall result and the range.

Why four runs

BFCL samples: vLLM raises BFCL's temperature of 0.001 to 0.01, and the model's default top_p/top_k apply. We ran Qwen3-0.6B twice with identical settings, once on simple_python,irrelevance and once on the 7 non-live categories. Between the two runs, the same arm changed its verdict on 2–13% of the cases per category, and the unweighted Δ moved from +0.83 to −0.42 on the first set and from −1.01 to +0.44 on the second. That leaves little margin under a 2-point tolerance. In four identical runs with this PR's setup (see Validation), the unweighted Δ ranged from −0.74 to +2.33. Greedy decoding does not remove the noise either: at temperature 0 the same arm still changed its verdict on about as many cases between runs, so it comes from batching concurrent requests, not from sampling. Averaging four runs halves the noise, and because each run needs only two of the node's eight GPUs, it adds no wall-clock time.

Switching the AMD legs on

Until then, the AMD legs run only from a workflow_dispatch that names one with only, and never on pull_request, so the runner only ever takes code from main. To switch them on:

  1. Register AMD's runner (8× MI325X) with the label 8-gpu-mi325x, or set the repository variable SMG_RUNNER_AMD_GPU_8 to its label.
  2. Set the repository variable SMG_RUN_AMD_LEGS to true.

The runner needs ROCm 7.2 installed, glibc 2.34 or newer, network access to PyPI, wheels.vllm.ai and Hugging Face, the two models under /models (or room to download them), and lsof for the GPU cleanup. The ROCm release matters because the wheels are built for 7.2 and load some of its system libraries; scripts/ci_install_vllm_rocm.sh checks it before installing. scripts/ci_setup_python_venv.sh provisions Python 3.12, which the ROCm wheels require.

Other changes

  • ci_killall_sglang.sh rocm nuke_gpus now kills the processes holding /dev/kfd, the ROCm counterpart of nuke_gpus. The CUDA path is unchanged.
  • The always() teardown step now sets BFCL_RUN_DIR. Without it, launch_arm.sh stop looked for pidfiles in /tmp/bfcl_ab instead of $RUNNER_TEMP/bfcl_run, so on every leg this backstop never found the servers it is meant to stop.
  • scripts/bfcl/README.md and CONTRIBUTING.md describe the AMD legs and the two variables.

Validation

  • run_ab.py --combine has four new unit tests; all 7 tests in test_bfcl_run_ab.py pass, and ruff and pre-commit pass.

  • We ran the matrix script for scheduled, dispatched and pull-request runs, with the AMD variables set and unset. The H100 and Blackwell legs come out exactly as before; the AMD legs appear only as described above.

  • On an MI355X node, with the same scripts an AMD runner would run (ci_setup_python_venv.sh, ci_install_vllm_rocm.sh, the SMG wheel and bfcl-eval, then run_repeats.sh). The first attempt, on a host whose system ROCm is 7.1.1, failed when torch imported amdsmi, which needs 7.2's libamd_smi; that is why the install script now checks the ROCm release first. The second ran with ROCm 7.2.3 libraries; the setup took about 2 minutes. Scores are vLLM / SMG, with Δ = SMG − vLLM:

    • gpt-oss-amd as configured, four runs of the 17 categories: the eight servers were up in about 8 minutes (loading from network storage) and scoring took 41 minutes. The mean of the four runs was 52.70 / 52.23 (−0.47) unweighted and 66.75 / 66.40 (−0.34) weighted, and the runs' unweighted Δ were −0.03, −0.11, −1.25 and −0.48. The 2026-10-05 H100 run, with the same vLLM 0.27.1, scored 52.73 / 52.00 and 67.10 / 66.22.
    • qwen3.8-amd's setup on simple_python,irrelevance, four runs: the mean was 86.44 / 86.12 (−0.31) unweighted and 87.97 / 87.73 (−0.23) weighted, with run Δ of −0.13, +0.04, −1.08 and −0.08 unweighted. On these two categories the H100 run scored 86.38 / 86.42 and 87.97 / 87.81. The slowest single case took about 20 minutes, which is why the leg keeps the H100 leg's 5-hour cap.
    • Noise, Qwen3-0.6B on the 7 non-live categories, four runs at each temperature: at BFCL's default, the runs' unweighted Δ were −0.39, −0.74, +2.33 and +0.29, and between two runs the same arm changed its verdict on 3.0–11.0% of the cases per category. With greedy decoding (temperature 0) it was 2.8–10.2%.

    The MI325X is gfx942 and the MI355X gfx950, so the first dispatch on the MI325X runner is the check for that hardware.

For reviewers

  1. AMD will provide and operate the runner. Registering it needs an org admin, with a registration token or a runner group limited to this repository. Does a repository-level runner with the label above work for you, or would you prefer another arrangement?
  2. The AMD legs gate on the mean of four runs, while the H100 legs run once. run_repeats.sh works for any leg whose arms leave room on its node; the H100 legs, at TP=2 on a 4-GPU runner, do not.
  3. AMD will follow up on failures of the AMD legs in the nightly-triage issue.

Add gpt-oss-amd and qwen3.8-amd to nightly-bfcl.yml: the H100 legs'
models and parsers on an AMD Instinct self-hosted runner, with vLLM
from the ROCm wheel of the CI pin (scripts/ci_install_vllm_rocm.sh).
Each arm fits one GPU, so a leg runs four A/B pairs side by side on
one node (arm_mode: repeated, scripts/bfcl/run_repeats.sh) and gates
on their mean (run_ab.py --combine). The legs stay off until the
runner is registered and SMG_RUN_AMD_LEGS is "true", and never run
on pull_request.

Also add a ROCm GPU cleanup to ci_killall_sglang.sh (rocm nuke_gpus),
and set BFCL_RUN_DIR for the always() teardown step, which looked for
pidfiles in the wrong directory.

Signed-off-by: Chun Fang <chun.fang@amd.com>
@coderabbitai

coderabbitai Bot commented Oct 6, 2026

Copy link
Copy Markdown

Important

Draft PR not reviewed

Draft PRs are not automatically reviewed by default.

  • Trigger a manual review

To automatically review draft PRs, update your CodeRabbit configuration:

reviews:
  auto_review:
    drafts: true
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Comment @coderabbitai help to get the list of available commands.

@github-actions github-actions Bot added documentation Improvements or additions to documentation ci CI/CD configuration changes tests Test changes labels Oct 6, 2026
inputs:
only:
description: "Run only this matrix leg (qwen3.8|gpt-oss|deepseek-v4.1|minimax-m3|glm-5.3-flash); empty = all"
description: "Run only this matrix leg (qwen3.8|gpt-oss|deepseek-v4.1|minimax-m3|glm-5.3-flash|gpt-oss-amd|qwen3.8-amd); empty = all"

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

should this be in here or add another variable to define hardware of nvidia or amd or both with default being both?

Comment on lines +126 to +128
# AMD legs: the AMD runner's label, and "true" once that runner is online.
AMD_RUNNER: ${{ vars.SMG_RUNNER_AMD_GPU_8 }}
AMD_ENABLED: ${{ vars.SMG_RUN_AMD_LEGS }}

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

has an stable fleet of AMD CI runners been hooked up into the upstream smg repo yet?

@@ -0,0 +1,101 @@
#!/usr/bin/env bash

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

what the need for this verus on existing NVIDIA? can we keep scope to just amd enablement

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ci CI/CD configuration changes documentation Improvements or additions to documentation tests Test changes

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants