Skip to content

fix(rocm): make hipBLASLt GEMM state safe across threads and streams #3911

fix(rocm): make hipBLASLt GEMM state safe across threads and streams

fix(rocm): make hipBLASLt GEMM state safe across threads and streams #3911

Workflow file for this run

# PR / push CI for lightweight quality gates. Runs cargo-deny (license /
# advisory) and cargo-fmt on every touched-Rust change so formatting drift and
# license issues are caught at PR time.
#
# The general unit suite is not gated at PR time, and clippy is gated only in
# the narrow shape described below. Both used to run on the self-hosted Apple
# Silicon runner, first here and then briefly in release.yml, and in both cases
# consumed ~30 min per run, blocking either PRs or releases on a shared
# resource for failures that `make verify` reliably catches on the developer's
# machine in a fraction of the time (#21, #23).
#
# The `clippy` job below is deliberately not a return to that arrangement. It
# runs `cargo clippy -p mlxcel --lib --tests -- -D warnings` at default
# features on the self-hosted GB10 runner with its own persistent target
# directory, so it neither touches the Apple Silicon runner nor pays a cold
# build after its first run. It exists because #916 merged an `err_expect`
# that reddened `make verify` for every contributor on every platform and sat
# on `main` until the nightly backstop caught it a day later (#1283): the
# per-crate lint that catches it is cheap, and only its absence at PR time was
# expensive.
#
# They are not absent from every workflow, though, so do not read the above as
# "nothing anywhere runs them":
#
# - pipeline-parallel-ci.yml runs `cargo clippy -p mlxcel --lib --tests
# -- -D warnings`, and reaches `cargo test` through scripts/ci/run-pp-*.sh
# with `distributed::`-prefixed selectors. It is debug profile on
# ubuntu-latest and path-filtered to src/distributed/pipeline/** and its
# siblings, so it never runs on a model-port PR and never selects a model
# or CLI test module.
# - nightly-verify.yml runs the full `make verify` (fmt + clippy + test) once
# a day on the self-hosted Apple Silicon runner and files an issue when it
# fails. It is a backstop against a red `main`, not a PR gate; see #939 for
# the two deterministic failures that sat on `main` before it existed, and
# for why it is nightly rather than per-PR.
#
# Quality gate still lives locally at PR time. Pre-push checklist:
# make verify # fmt + clippy(workspace, metal,accelerate, -D warnings)
# # + test(workspace, test-fast)
# make verify-clean # same, after `cargo clean` — use when clippy's
# # per-crate cache may be hiding a regression
#
# Note on CUDA gating: running the CUDA test suite stays exclusive to
# release.yml because it requires a Linux self-hosted runner (currently only the
# GB10 node used for release builds). Adding a PR-level CUDA *test* gate would
# double runner cost for limited additional safety on PRs that don't touch
# CUDA-specific code paths. Compiling CUDA is a different question and four
# jobs below do it: `xla-compile` and `xla-link` at the shipped sm_121,
# `cuda-blockfloat` at the same list, and `cuda-sm70-compile` at sm_70, each on a
# narrow path filter. `cuda-blockfloat` is the single exception to the sentence
# above, and a narrow one: it runs the block-float quantization tests and only
# those, because until issue #1934 no build anywhere compiled MLX's hardware
# NVFP4/MXFP4 converter, so those tests had never once exercised the path they
# are named for.
name: CI
on:
push:
branches: [main]
pull_request:
branches: [main]
concurrency:
# One CI run per branch, because seven jobs below land on GB10, the single
# self-hosted Linux runner that also serves the release build: `clippy`,
# `cpu-link`, `webui-installed-artifact`, `xla-compile`, `xla-link`,
# `cuda-sm70-compile`, and `cuda-blockfloat`. Without a group, a run
# for a commit nobody is waiting on any more does not merely sit idle, it
# holds a place in line ahead of the runs that matter. Measured over
# 2026-09-10 KST: 113 runs created, 14 open at once, 11 of those stacked on
# four branches, and six cancelled by hand purely to free the runner (#1774).
#
# `main` is deliberately excluded from the cancellation, which is what the
# `cancel-in-progress` expression encodes. A push-to-main run is the one that
# catches what a stale PR base hid, and it is the last gate before a release
# builds from that commit; killing it because the next merge landed 90 seconds
# later discards the only signal nothing else produces. The queueing cost of
# that choice is bounded rather than unbounded, because cancelling a *pending*
# run is inherent to grouping and `cancel-in-progress` governs only the
# *running* one: GitHub keeps the in-flight `main` run plus at most one
# pending run behind it, so back-to-back merges queue two deep and not N deep.
# What that leaves in place, stated plainly rather than discovered later, is
# that an intermediate `main` commit can have its run cancelled while still
# pending; whatever commit is at the head of `main` is always the newest in
# the group and so always gets a run, which is the property that matters
# before a release is cut from it. nightly-verify.yml made the same trade for
# the same reason (its group is unconditional, having no PR half to key on).
#
# `github.event.pull_request.number || github.ref` is the shape already used
# by pipeline-parallel-ci.yml and by the two per-job groups below, and it
# keeps a PR run and a `main` push run in distinct groups. The per-job groups
# in `xla-link` and `cuda-sm70-compile` stay as they are: workflow-level and
# job-level groups compose, and those two are keyed differently on purpose.
group: ci-${{ github.workflow }}-${{ github.event.pull_request.number || github.ref }}
cancel-in-progress: ${{ github.ref != 'refs/heads/main' }}
permissions:
contents: read
jobs:
changes:
name: Detect changes
runs-on: ubuntu-latest
permissions:
contents: read
pull-requests: read
outputs:
rust: ${{ steps.filter.outputs.rust }}
mlx_pin: ${{ steps.filter.outputs.mlx_pin }}
xla_link: ${{ steps.filter.outputs.xla_link }}
cuda_arch: ${{ steps.filter.outputs.cuda_arch }}
rocm: ${{ steps.filter.outputs.rocm }}
cpu_link: ${{ steps.filter.outputs.cpu_link }}
workflows: ${{ steps.filter.outputs.workflows }}
webui_bundle: ${{ steps.filter.outputs.webui_bundle }}
steps:
- uses: actions/checkout@v7
with:
persist-credentials: false
- uses: dorny/paths-filter@v4
id: filter
with:
# Required for the `!` patterns below to mean anything. Under the
# default `some`, a negated pattern is just another pattern to match,
# so `'src/lib/mlx-cpp/**'` plus `'!src/lib/mlx-cpp/patches-rocm/**'`
# still matches a ROCm-only change: the action's own README says the
# exclusion syntax is ignored "UNLESS you specify 'every' or
# 'some-with-excludes'". Safe to set globally because no filter here
# used negation before this change, and for a filter with no negated
# rule `some-with-excludes` reduces to `some` (lablup/mlxcel#1992).
predicate-quantifier: 'some-with-excludes'
filters: |
rust:
- '**/*.rs'
- '**/Cargo.toml'
- 'Cargo.lock'
- 'deny.toml'
- 'build.rs'
- 'scripts/iree/**'
mlx_pin:
- 'src/lib/mlx-cpp/CMakeLists.txt'
- 'src/lib/mlxcel-core/build_support/mlx_pin.rs'
- 'src/lib/mlxcel-mlx-pin/**'
- 'scripts/ci/mlx_pinned_commit.sh'
- 'scripts/ci/mlx_pinned_commit_test.sh'
# Narrower than `rust` above: only the paths that can move the IREE
# link recipe itself, so the release link job below does not run on
# every Rust PR. See the `xla-link` job comment for the rationale.
xla_link:
- 'build.rs'
- 'src/lib/mlxcel-xla/build.rs'
- 'src/lib/mlxcel-xla/csrc/**'
- 'scripts/iree/**'
- 'rust-toolchain.toml'
# The CUDA half of the same link line: mlxcel-core's build
# script names the CUDA libraries, and the MLX pin decides which
# of them libmlx.a needs. See the `xla-link` job comment.
- 'src/lib/mlxcel-core/build.rs'
- 'src/lib/mlx-cpp/CMakeLists.txt'
# Everything that can compile differently per CUDA architecture: the
# MLX overlays and turbo kernels, the pinned MLX commit, and the
# build scripts that hand CMake the architecture list. Rust sources
# outside these paths cannot break an arch-conditional compile, so
# the sm_70 gate below does not run on every Rust PR. See the
# `cuda-sm70-compile` job comment for why sm_70 specifically.
cuda_arch:
- 'src/lib/mlx-cpp/**'
# Two subtrees under that glob cannot change what nvcc compiles,
# and these are the heaviest jobs on the shared GB10 runner. Most
# of epic #1801's pull requests touch only `patches-rocm/`, so
# before this exclusion every one of them woke both CUDA jobs
# (lablup/mlxcel#1992).
#
# Note what is NOT excluded: `patches/` as a whole. Despite the
# name it is not the Metal overlay. `patches/mlx/backend/cuda/`
# holds 24 CUDA sources, applied and compiled for CUDA targets by
# `src/lib/mlx-cpp/CMakeLists.txt:59-69`, against 4 Metal files in
# `patches/mlx/backend/metal/`. Excluding `patches/**` would drop
# real CUDA coverage, so only the Metal subtree is named.
- '!src/lib/mlx-cpp/patches-rocm/**'
- '!src/lib/mlx-cpp/patches/mlx/backend/metal/**'
- 'src/lib/mlxcel-core/build.rs'
- 'src/lib/mlxcel-core/build_support/**'
- 'build.rs'
# Everything that can change what the ROCm build compiles, links or
# runs: the HIP overlay, the shared MLX build inputs, the bridge
# sources the overlay is called through, and the pinned MLX commit.
# A pin bump must trigger this, because the ROCm half of every bump
# now lives in `patches-rocm/` and nothing else in CI links it.
rocm:
- 'src/lib/mlx-cpp/**'
# The mirror of the exclusion above: neither the Metal files nor
# the CUDA overlays reach a HIP compile. `patches/` is split by
# subtree here for the same reason it is there, that the directory
# holds both backends' files.
- '!src/lib/mlx-cpp/patches/mlx/backend/metal/**'
- '!src/lib/mlx-cpp/patches/mlx/backend/cuda/**'
- '!src/lib/mlx-cpp/patches-cuda/**'
- 'src/lib/mlxcel-core/build.rs'
- 'src/lib/mlxcel-core/build_support/**'
- 'src/lib/mlxcel-core/cpp/**'
- 'build.rs'
- '**/Cargo.toml'
- 'Cargo.lock'
# What can break the CPU-only link of `mlxcel-core`: the Rust
# sources and manifests, plus the bridge C++ that `rust` above does
# not match. The bug that motivated the `cpu-link` job lived in
# `src/lib/mlx-cpp/turbo/` (#2108). No overlay excludes: that job
# costs seconds warm, unlike the CUDA jobs those excludes protect.
cpu_link:
- '**/*.rs'
- '**/Cargo.toml'
- 'Cargo.lock'
- 'build.rs'
- 'rust-toolchain.toml'
- 'src/lib/mlx-cpp/**'
- 'src/lib/mlxcel-core/cpp/**'
- 'src/lib/mlxcel-core/build_support/**'
# The workflow files themselves. Deliberately NOT folded into the
# filters above: those gate what gets compiled, and a workflow edit
# compiles nothing. What a workflow edit needs is the lint job and
# the advisory below, both hosted.
workflows:
- '.github/workflows/**'
- '.github/actionlint.yaml'
webui_bundle:
- 'webui/**'
- 'src/webui/assets/**'
- 'src/server/**'
- 'src/cli/**'
- 'src/bin/mlx_server.rs'
- 'src/main.rs'
- 'tests/fixtures/webui/**'
- 'scripts/webui/**'
- 'scripts/ci/*webui*'
- 'Cargo.toml'
- 'Cargo.lock'
- 'Makefile'
deny:
name: cargo-deny
needs: changes
if: needs.changes.outputs.rust == 'true'
runs-on: ubuntu-latest
permissions:
contents: read
steps:
- uses: actions/checkout@v7
with:
persist-credentials: false
- uses: EmbarkStudios/cargo-deny-action@v2
with:
command: check
fmt:
name: cargo-fmt
needs: changes
if: needs.changes.outputs.rust == 'true'
runs-on: ubuntu-latest
permissions:
contents: read
steps:
- uses: actions/checkout@v7
with:
persist-credentials: false
# The tag on this action IS the toolchain version it installs, so it must
# track `rust-toolchain.toml`'s `channel` rather than be bumped on its own.
# Dependabot reads the tag as an ordinary version and will try to raise it;
# `.github/dependabot.yml` ignores this action for that reason.
- uses: dtolnay/rust-toolchain@1.97.1
with:
components: rustfmt
- run: cargo fmt --all -- --check
clippy:
name: cargo-clippy
needs: changes
# On the repository guard, because it is easy to read more into it than it
# does. It keeps this job from queueing forever in a *fork of* this
# repository, which has no GB10 runner of its own. It does not stop a pull
# request opened *from* a fork against lablup/mlxcel: on that event
# `github.repository` is the base repository, so the guard is true and the
# job runs here, executing the PR's `build.rs` and `scripts/iree/**` as the
# runner's own user. What actually gates that case is the repository's
# Actions fork-PR approval policy, currently `first_time_contributors`.
# The same reading applies to `xla-compile` and `xla-link` below.
if: >-
github.repository == 'lablup/mlxcel'
&& (needs.changes.outputs.rust == 'true' || contains(github.event.pull_request.labels.*.name, 'ci:full'))
runs-on: GB10
permissions:
contents: read
timeout-minutes: 60
steps:
- uses: actions/checkout@v7
with:
persist-credentials: false
- name: Use a persistent target directory
run: |
# Clippy emits different artifacts than a build, so it gets its own
# directory rather than sharing (and repeatedly invalidating) the
# xla-compile or release caches. Cold on first run, warm after.
CARGO_TARGET="$HOME/.cargo-target/mlxcel-clippy-ci"
mkdir -p "$CARGO_TARGET"
echo "CARGO_TARGET_DIR=$CARGO_TARGET" >> $GITHUB_ENV
# Default features (`surgery`, no accelerator). The lint that motivated
# this job (#1283) reproduces on every feature set, so the cheapest one
# that compiles the crate is enough, and it is the same command
# pipeline-parallel-ci.yml already runs on its own path filter.
#
# On the runner choice, measured rather than assumed, so the question does
# not get re-litigated from scratch. This job is here and not on
# `ubuntu-latest` because `src/lib/mlxcel-core/build.rs` builds MLX through
# cmake unconditionally: the accelerator features pick a backend, they do
# not decide whether the C++ builds. A GitHub-hosted runner would pay that
# cold on every run, since the one `ubuntu-latest` job this repository
# already has (`pipeline-parallel-ci.yml`) caches the cargo registry and
# not `target/`, and it takes about 28 minutes. Here the same command is
# 2m44s cold and 27 to 40 seconds warm off the persistent target directory.
#
# The cost of staying on GB10 is that this runner also serves
# `xla-compile` and the release build, so a release in flight queues this
# job behind it and delays a merge. That happened once on 2026-08-22
# (#1301). It is a delay and not a failure: a queued job consumes nothing,
# and the release was not slowed by it. Trading a rare queue wait for
# roughly 20 minutes on every PR would re-create the cost objection that
# removed clippy from PR-time CI in the first place (#21, #23).
- name: Clippy
run: |
cargo clippy \
-p mlxcel \
--lib \
--tests \
-- -D warnings
cpu-link:
name: CPU-only link
needs: changes
# Links the `mlxcel-core` test binary with no GPU feature. A Linux build
# without one compiles MLX with no GPU backend, so bridge C++ that calls a
# GPU-only MLX helper without the `MLXCEL_BRIDGE_GPU_BACKEND` guard fails
# only here, at link time: `cargo check` does not link, and every other
# GB10 job passes `--features cuda` (#2108). `--no-run` links without
# running the tests. The repository guard reads as on `clippy` above.
if: >-
github.repository == 'lablup/mlxcel'
&& (needs.changes.outputs.cpu_link == 'true' || contains(github.event.pull_request.labels.*.name, 'ci:full'))
runs-on: GB10
permissions:
contents: read
timeout-minutes: 30
steps:
- uses: actions/checkout@v7
with:
persist-credentials: false
- name: Use a persistent target directory
run: |
# Its own directory: a CPU-only MLX tree would otherwise evict the
# CUDA build in any directory it shared. Measured on GB10 for #2108,
# about 2.5 minutes cold (including the CPU-only MLX tree) and about
# 6 seconds to relink after a one-file change.
CARGO_TARGET="$HOME/.cargo-target/mlxcel-cpu-link-ci"
mkdir -p "$CARGO_TARGET"
echo "CARGO_TARGET_DIR=$CARGO_TARGET" >> "$GITHUB_ENV"
- name: Link the CPU-only mlxcel-core test binary
run: cargo test -p mlxcel-core --profile test-fast --lib --no-run
crate-versions:
name: crate versions
# Deliberately not behind the `changes` filter. It needs no toolchain and
# runs in seconds, and a version-consistency gate that can be skipped is
# exactly the gate that gets skipped on the one PR that breaks it.
runs-on: ubuntu-latest
permissions:
contents: read
steps:
- uses: actions/checkout@v7
with:
persist-credentials: false
# Gating, not advisory. From v0.3.3 onward the root `mlxcel` and every
# member crate that ships with it carry the same `[package] version`, and
# a release has to bump several manifests at once. The list used to live
# in prose (CLAUDE.md, the release skill) and went stale the moment a
# member was added: `mlxcel-mlx-pin` tracked the root version from the day
# it was created without either document saying so. The script inverts the
# rule, so a new member fails this job until it is either bumped with the
# root or declared independent with a reason. No toolchain needed.
- name: Check every version-tracking crate carries the root version
run: python3 scripts/ci/check_crate_versions.py
kernel-dtype-keys:
name: kernel dtype keys
# Deliberately not behind the `changes` filter, for the same reason as
# `crate-versions`: it needs no toolchain, runs in seconds, and the defect it
# guards against is silent. MLX's CUDA JIT memoises a compiled kernel under a
# name that does not carry the input dtypes, so a launch whose template args
# are all ints serves the first dtype's compiled module to every later dtype
# at the same geometry, reading buffers through the wrong pointer type.
# Issues #1053 and #1054 are two symptoms of exactly that. macOS never sees
# it, because Metal's key does carry the dtypes, so this job is the only
# place the omission is caught without CUDA hardware. ROCm's
# `fast::hip_kernel` keys its cache the same way, so HIP launches are held
# to the same rule (#1875). The second step is the checker's own negative
# coverage: it moves launches out of scope in a throwaway copy and asserts
# the checker fails, so a refactor cannot empty the scope silently.
runs-on: ubuntu-latest
permissions:
contents: read
steps:
- uses: actions/checkout@v7
with:
persist-credentials: false
- name: Check every CUDA and HIP JIT launch keys its cache on the input dtypes
run: python3 scripts/ci/check_kernel_dtype_keys.py
- name: Check the dtype-key checker rejects a shrunk or moved scope
run: bash scripts/ci/check_kernel_dtype_keys_test.sh
kernel-port-dispatch:
name: kernel port dispatch
# Not behind the `changes` filter, for the same reasons as the job above: no
# toolchain, seconds to run, and the defect it guards against is silent.
#
# The defect is specifically a convention decaying. Choosing a kernel's port
# with `use_cuda ? cuda : metal` reads "not CUDA" as "Metal", which was true
# with two backends and wrong with three: on ROCm the false arm threw and,
# because the bridge declarations were not `Result`, the throw crossed a
# `noexcept` extern into `std::terminate`. #1803 fixed the idiom and left
# nine sites shaped that way; #1885 and #2018 then guarded them by hand, one
# launcher at a time, months apart, each after it had already aborted a gate
# run, and the hand-written guards drifted in wording and twice named a
# predicate that did not exist. A convention that must be re-applied at every
# new launcher gets missed at some of them, and only the backend that lacks
# the port ever finds out. This job is what keeps the single helper single.
runs-on: ubuntu-latest
permissions:
contents: read
steps:
- uses: actions/checkout@v7
with:
persist-credentials: false
- name: Check every fused-kernel launcher routes through select_kernel_port
run: python3 scripts/ci/check_kernel_port_dispatch.py
- name: Check the port-dispatch checker rejects an unvalidated wave64 port
run: bash scripts/ci/check_kernel_port_dispatch_test.sh
binary-assets:
name: binary assets
# Deliberately not behind the `changes` filter, for the same reason as
# `crate-versions` and `kernel-dtype-keys`: it needs no toolchain, runs in
# seconds, and the defect it guards against is silent. Binary blobs never
# delta-compress and deleting one later does not shrink the history, so the
# cost of adding one only shows up as clone time long after the commit.
# `webui/tests/screenshots/` reached 24 PNGs and ~1.9 MB without any review
# step noticing, because each commit looked reasonable on its own and
# nothing looked at the aggregate. The check inverts the rule: an
# undeclared binary path fails until someone declares it with a reason.
runs-on: ubuntu-latest
permissions:
contents: read
steps:
- uses: actions/checkout@v7
with:
persist-credentials: false
- name: Check every tracked binary file is declared with a size budget
run: python3 scripts/ci/check_binary_assets.py
- name: Check the guard still fails on undeclared and oversize binaries
run: bash scripts/ci/check_binary_assets_test.sh
cuda-arch-lists:
name: CUDA architecture lists
# Deliberately not behind the `changes` filter, like `crate-versions` and
# `binary-assets`: it needs no toolchain, runs in seconds, and both defects
# it guards against are silent.
#
# MLX compiles its hardware NVFP4/MXFP4 converters
# (`cvt.rn.satfinite.e2m1x2.f32` in
# mlx/backend/cuda/quantized/nvfp4_quantize.cuh) only under an
# architecture-specific target, and the lists here are plain, so what
# supplies them is the per-source injection in
# `src/lib/mlx-cpp/CMakeLists.txt`. Remove that and the plain lists quietly
# go back to compiling the converters out of every artifact and every test,
# which is where issue #1934 started, so this asserts it is still there.
#
# The other direction is the obvious fix: putting `121a` in a list. That
# compiles every translation unit architecture-specific for a converter one
# of them reaches, costs 12.7 MB of archive and replaces the machine code of
# every decode kernel on every Blackwell host. Two tokens in a YAML `env`
# block, and no build output names the difference, so it is rejected here by
# name. So is the `f` spelling, which reads as a milder `a` and does not
# compile at all.
#
# The `cuda-blockfloat` job below asserts the same two properties on the
# compiled artifact, which is stronger, but it costs a CUDA build and is
# path-filtered. This one runs on every PR.
runs-on: ubuntu-latest
permissions:
contents: read
steps:
- uses: actions/checkout@v7
with:
persist-credentials: false
- name: Every Blackwell architecture list carries a hardware-capable entry
run: python3 scripts/ci/check_cuda_arch_lists.py
- name: Check the guard still fails on plain-only and family-specific lists
run: bash scripts/ci/check_cuda_arch_lists_test.sh
license-headers:
name: license headers
# Deliberately not behind the `changes` filter, like `crate-versions` and
# `kernel-dtype-keys`: it needs no toolchain and runs in seconds.
#
# This gate exists because coverage without enforcement did nothing. Issue
# #1707 found `benches/audio_fft.rs` shipping unlicensed and traced it to
# `benches` missing from the script's `TARGET_ROOTS`. That was only the
# visible instance: the roots list already named `src`, `examples` and
# `tests`, and 95 files under those three roots were unlicensed anyway,
# because nothing ever ran the script. Fixing the list would not have
# caught a single one of them. Running it in CI does.
#
# `--check` writes nothing and reports the files the default insert mode
# would touch, reusing the same `has_existing_header` predicate the writer
# branches on, so the gate can never demand an edit that running
# `python3 scripts/insert_apache_header.py` does not produce. Files that
# already carry someone else's provenance (the vendored MLX patches under
# src/lib/mlx-cpp/patches/ keep Apple's copyright) are skipped by that same
# predicate and must stay that way: this gate must never stamp a Lablup
# header onto upstream-derived code.
#
# The second step is the script's own unit suite, which nothing else runs:
# tests/test_insert_apache_header.py sits outside python/, and
# python.yml's pytest step is scoped to `python/tests` behind a
# `python/**` path filter.
runs-on: ubuntu-latest
permissions:
contents: read
steps:
- uses: actions/checkout@v7
with:
persist-credentials: false
- name: Every owned source file carries the license header
run: python3 scripts/insert_apache_header.py --check
- name: Header script unit tests
run: python3 -m unittest discover -s tests -p 'test_insert_apache_header.py' -v
llama-compat-manifest:
name: llama-compat manifest
# Deliberately not behind the `changes` filter, like `crate-versions`: it
# needs no toolchain and runs in seconds. Validates the checked-in
# llama-server b10621 compatibility manifest (compat/llama-server/b10621/,
# issue #1443, epic #1431): pinned counts (249 help entries, 323 long
# spellings, 134 LLAMA_* env vars), one policy state per entry (five
# states since #1499: supported / aliased / not_applicable / deferred /
# by_design), issue links on deferred entries, test ids on non-supported
# entries, the entry key allowlist, the rule that a non-empty
# `divergence` forbids `supported`, the by_design obligations (non-empty
# divergence, notes, a resolving test pointer, and a well-formed
# `rationale` whose policy kind must state its revisit condition, with
# `rationale` null on every other state), and canonical serialization.
# --check-issues-open additionally asserts every issue a deferred entry
# links is still open, so closing a chain issue without flipping its
# manifest entries fails here; by_design entries never feed that check,
# because a permanent, tested divergence has no issue left to keep open. The second step
# is the validator's own negative coverage: it mutates a throwaway copy of
# the manifest and asserts the gate rejects it. The binary-facing half of the
# gate (option spellings, env bindings, defaults against both server
# binaries, hidden arguments included, plus mounted routes and native
# request fields) runs inside the workspace test suite as
# tests/llama_compat_manifest.rs and src/server/llama_compat_tests.rs.
runs-on: ubuntu-latest
permissions:
contents: read
issues: read
env:
GH_TOKEN: ${{ github.token }}
steps:
- uses: actions/checkout@v7
with:
persist-credentials: false
- name: Validate the b10621 compatibility manifest
run: python3 scripts/ci/check_llama_compat_manifest.py --check-issues-open
- name: Validator negative coverage
run: bash scripts/ci/check_llama_compat_manifest_test.sh
webui-bundle:
name: WebUI bundle
needs: changes
if: needs.changes.outputs.webui_bundle == 'true'
runs-on: ubuntu-24.04
permissions:
contents: read
steps:
- uses: actions/checkout@v7
with:
persist-credentials: false
- uses: actions/setup-node@v7
with:
node-version: '26.5.1'
- name: Install pnpm
run: npm install --global pnpm@11.18.0
- name: Install WebUI dependencies
run: pnpm --dir webui install --frozen-lockfile
- name: Typecheck WebUI
run: pnpm --dir webui run typecheck
- name: Lint WebUI
run: pnpm --dir webui run lint
# Issue #1903: data-theme is always <family>-<scheme>, so a selector keyed
# on a bare mode (or on the host OS scheme) silently stops applying. The
# gate evaluates every theme selector against the ids theme.ts ships.
- name: Check WebUI theme selectors
run: pnpm --dir webui run check:theme-selectors
- name: Test WebUI scaffold
run: pnpm --dir webui run unit
- name: Install browser runtime
run: pnpm --dir webui exec playwright install --with-deps chromium firefox webkit
# Kept after the pixel baselines were dropped. The browser suite no longer
# compares images, but expectSafeLayout, expectCompactToolbarHitTargets and
# expectTextScaleLabelsReachable all assert measured geometry, and the CJK
# drawer variant needs a Korean face to lay out at all. An unpinned runner
# font set moves those measurements, so the pin is now a layout-stability
# dependency rather than a rendering one.
- name: Install WebUI layout fonts
run: |
sudo apt-get install -y --no-install-recommends fonts-dejavu-core=2.37-8 fonts-dejavu-extra=2.37-8 fonts-wqy-zenhei=0.9.45-8
fc-cache -f
- name: Log WebUI browser font environment
run: |
uname -a
cat /etc/os-release
node --version
pnpm --version
pnpm --dir webui exec playwright --version
pnpm --dir webui exec playwright install --dry-run chromium firefox webkit
fc-match sans-serif || true
fc-match system-ui || true
fc-match 'system-ui:weight=bold' || true
fc-match Arial || true
fc-list | wc -l
- name: Run WebUI Chromium layout and accessibility scaffold
env:
MLXCEL_WEBUI_FONT_DIAGNOSTICS: '1'
run: pnpm --dir webui run browser
- name: Ensure headed WebUI browser display helper
run: |
if ! command -v xvfb-run >/dev/null 2>&1; then
sudo apt-get install -y --no-install-recommends xvfb
fi
command -v xvfb-run
- name: Run WebUI browser engine smoke on Chromium, Firefox, and WebKit
run: xvfb-run -a --server-args='-screen 0 1920x1080x24' pnpm --dir webui run browser:all
- name: Upload WebUI browser failure artifacts
if: failure()
uses: actions/upload-artifact@v7
with:
name: webui-browser-failure-artifacts-${{ runner.os }}
path: |
webui/test-results/**/*.png
if-no-files-found: warn
- name: Verify generated WebUI bundle
run: python3 scripts/webui/build_bundle.py --verify
webui-contract:
name: WebUI contract
# Contract-first gate for the bundled WebUI epic. It needs no Rust
# toolchain and intentionally runs as its own drift check: schema changes
# must update generated TypeScript DTOs and representative fixtures in the
# same PR before backend or frontend owners consume the contract.
runs-on: ubuntu-latest
permissions:
contents: read
steps:
- uses: actions/checkout@v7
with:
persist-credentials: false
- name: Install pinned contract verifier
run: python3 -m pip install -r scripts/ci/webui_contract_requirements.txt
- name: Validate WebUI schema, DTOs, and fixtures
run: python3 scripts/ci/check_webui_contract.py --self-test
webui-installed-artifact:
name: WebUI installed artifact
needs: changes
if: >-
github.repository == 'lablup/mlxcel'
&& (needs.changes.outputs.webui_bundle == 'true'
|| needs.changes.outputs.rust == 'true'
|| contains(github.event.pull_request.labels.*.name, 'ci:full'))
runs-on: GB10
permissions:
contents: read
# Two cold builds plus the Activity gate step's 75 minutes (#1949).
timeout-minutes: 120
env:
# What releases build for this runner's architecture, pinned rather than
# inferred so the artifact this job verifies carries the same device code
# the published one does whatever GPU the runner turns out to have. Since
# issue #1943 auto-detection produces this same `121` on a GB10, so the pin
# no longer corrects a shape mismatch: it states the list instead of
# inferring it, which is what makes the job reviewable and what keeps it
# right if this ever runs on something other than a GB10.
MLX_CUDA_ARCHITECTURES: "121"
# `test-fast` sets `incremental = true`, and this job's persistent target directory once kept
# incremental state that linked with undefined `serde_json` generics from one codegen unit
# on five consecutive attempts until that state was cleared by hand (run 36988516810,
# #2111). This job builds each artifact once per run and gains little from incremental
# reuse, so it opts out. Cleaning and retrying on a link failure was rejected: it hides the
# cause and pays a full rebuild. Other GB10 target directories keep the profile default.
CARGO_INCREMENTAL: "0"
steps:
- uses: actions/checkout@v7
with:
persist-credentials: false
- name: Use a persistent target directory
run: |
CARGO_TARGET="$HOME/.cargo-target/mlxcel-webui-installed-ci"
mkdir -p "$CARGO_TARGET"
echo "CARGO_TARGET_DIR=$CARGO_TARGET" >> $GITHUB_ENV
- name: Build installed server and CLI artifacts once
run: cargo build --profile test-fast --features cuda,webui --bin mlxcel-server --bin mlxcel
- name: Preserve WebUI-enabled installed artifacts before feature-off build
run: |
mkdir -p "$RUNNER_TEMP/mlxcel-webui-installed"
cp "$CARGO_TARGET_DIR/test-fast/mlxcel-server" "$RUNNER_TEMP/mlxcel-webui-installed/mlxcel-server-webui"
cp "$CARGO_TARGET_DIR/test-fast/mlxcel" "$RUNNER_TEMP/mlxcel-webui-installed/mlxcel-webui"
- name: Build feature-disabled server and CLI artifacts once
run: cargo build --profile test-fast --no-default-features --features cuda --bin mlxcel-server --bin mlxcel
- name: Preserve feature-disabled installed artifacts
run: |
cp "$CARGO_TARGET_DIR/test-fast/mlxcel-server" "$RUNNER_TEMP/mlxcel-webui-installed/mlxcel-server-feature-off"
cp "$CARGO_TARGET_DIR/test-fast/mlxcel" "$RUNNER_TEMP/mlxcel-webui-installed/mlxcel-feature-off"
- uses: actions/setup-node@v7
with:
node-version: '26.5.1'
- name: Install pnpm for real-router browser harness
run: npm install --global pnpm@11.18.0
- name: Install WebUI dependencies for real-router browser harness
run: pnpm --dir webui install --frozen-lockfile
- name: Install Chromium runtime for real-router browser harness
run: pnpm --dir webui exec playwright install chromium
- name: Probe Chromium runtime for real-router browser harness
timeout-minutes: 1
run: |
set -euo pipefail
timeout 45s pnpm --dir webui exec node -e '
const { chromium } = require("@playwright/test");
(async () => {
let browser;
try {
browser = await chromium.launch();
const page = await browser.newPage({ viewport: { width: 800, height: 600 } });
await page.setContent("<main id=\"probe\">ready</main>");
const box = await page.locator("#probe").boundingBox();
if (!box || box.width <= 0 || box.height <= 0) {
throw new Error("Chromium geometry probe failed");
}
} finally {
if (browser) {
await browser.close();
}
}
})().catch((error) => {
console.error(`Chromium runtime probe failed: ${error.message}`);
process.exit(1);
});
'
- name: Run WebUI real-router browser harness
env:
MLXCEL_WEBUI_ROUTER_ARTIFACTS: webui/test-results/rust-router-ci
run: cargo test --profile test-fast --features cuda,webui --lib server::router_server::router_webui_playwright_harness_tests::real_router_browser_harness -- --ignored --exact --nocapture
- name: Unit-test the installed-artifact verifier helpers
run: |
python3 scripts/webui/measure_startup_tests.py
python3 scripts/webui/verify_installed_artifact_tests.py
python3 scripts/webui/verify_generated_key_helper_tests.py
python3 scripts/webui/verify_reverse_proxy_tests.py
python3 scripts/webui/verify_single_model_tests.py
python3 scripts/webui/verify_activity_performance_tests.py
python3 scripts/webui/summarize_activity_evidence_tests.py
python3 scripts/webui/activity_gate_host_tests.py
- name: Verify the installed Rust artifact serves the bundled WebUI
env:
WEBUI_INSTALLED_EVIDENCE: ${{ runner.temp }}/webui-installed-evidence.json
WEBUI_NETWORK_SANDBOX: require
WEBUI_REQUIRE_TLS: '1'
WEBUI_BUILD_SOURCE_HEAD: ${{ github.sha }}
WEBUI_FEATURE_OFF_BUILD_SOURCE_HEAD: ${{ github.sha }}
run: |
WEBUI_SERVER_BIN="$RUNNER_TEMP/mlxcel-webui-installed/mlxcel-server-webui" \
WEBUI_CLI_BIN="$RUNNER_TEMP/mlxcel-webui-installed/mlxcel-webui" \
WEBUI_FEATURE_OFF_SERVER_BIN="$RUNNER_TEMP/mlxcel-webui-installed/mlxcel-server-feature-off" \
WEBUI_FEATURE_OFF_CLI_BIN="$RUNNER_TEMP/mlxcel-webui-installed/mlxcel-feature-off" \
python3 scripts/webui/verify_installed_artifact.py
- name: Verify loopback reverse-proxy WebUI edge behavior
env:
WEBUI_REVERSE_PROXY_EVIDENCE: ${{ runner.temp }}/webui-reverse-proxy-evidence.json
run: |
python3 scripts/webui/verify_reverse_proxy.py \
--server-bin "$RUNNER_TEMP/mlxcel-webui-installed/mlxcel-server-webui" \
--cli-bin "$RUNNER_TEMP/mlxcel-webui-installed/mlxcel-webui" \
--build-source-head "${{ github.sha }}" \
--evidence "$WEBUI_REVERSE_PROXY_EVIDENCE"
- name: Check native-hidden Activity display prerequisites
run: |
set -euo pipefail
missing=0
for tool in Xvfb openbox flock nvidia-smi; do
if ! command -v "$tool" >/dev/null 2>&1; then
echo "::error::$tool is required on the GB10 runner for the native-hidden Activity gate"
missing=1
fi
done
if [ "$missing" -ne 0 ]; then
echo "::error::Provision Xvfb, openbox, flock, and nvidia-smi on the GB10 runner; this job does not use sudo or skip the native-hidden gate"
exit 1
fi
# The real-router harness runs headless, but the native-hidden Activity gate below needs a
# headed browser with a real window. Those fail differently: a headed launch with no window
# server reports a zero-sized content area rather than an error, which the gate only
# discovers after starting a server and loading a checkpoint. Probe it here, where the
# message names the cause.
- name: Probe headed Chromium geometry under a virtual display
timeout-minutes: 2
run: |
set -euo pipefail
timeout 90s xvfb-run -a --server-args='-screen 0 1600x1100x24' pnpm --dir webui exec node -e '
const { chromium } = require("@playwright/test");
(async () => {
let browser;
try {
browser = await chromium.launch({ headless: false });
const page = await browser.newPage({ viewport: { width: 700, height: 900 } });
await page.goto("data:text/html,<main>ready</main>");
const geometry = await page.evaluate(() => ({ innerWidth: window.innerWidth, innerHeight: window.innerHeight }));
console.log(`headed geometry: ${JSON.stringify(geometry)}`);
if (!(geometry.innerWidth > 0 && geometry.innerHeight > 0)) {
throw new Error("headed Chromium reported a zero-sized content area; the virtual display or window manager is not usable");
}
} finally {
if (browser) {
await browser.close();
}
}
})().catch((error) => {
console.error(`Headed Chromium probe failed: ${error.message}`);
process.exit(1);
});
'
- name: Cache Activity checkpoint fixture
id: activity-fixture-cache
uses: actions/cache@v6
with:
path: ${{ runner.temp }}/mlxcel-webui-activity-fixtures/qwen3-0.6b-4bit
key: ${{ runner.os }}-webui-activity-qwen3-0.6b-4bit-v1
- name: Download Activity checkpoint fixture
if: steps.activity-fixture-cache.outputs.cache-hit != 'true'
env:
GH_TOKEN: ${{ secrets.GITHUB_TOKEN }}
CI_FIXTURE_TAG: ci-fixtures/qwen3-0.6b-4bit-v1
CI_FIXTURE_ASSET: qwen3-0.6b-4bit.tar.gz
run: |
set -euo pipefail
fixture_root="$RUNNER_TEMP/mlxcel-webui-activity-fixtures"
mkdir -p "$fixture_root"
test ! -e "$fixture_root/qwen3-0.6b-4bit"
tmp_dir="$(mktemp -d)"
gh release download "$CI_FIXTURE_TAG" \
--pattern "$CI_FIXTURE_ASSET" \
--dir "$tmp_dir"
tar -xzf "$tmp_dir/$CI_FIXTURE_ASSET" -C "$fixture_root"
rm -rf "$tmp_dir"
test -s "$fixture_root/qwen3-0.6b-4bit/config.json"
# Deferred on 2026-09-16: the native hidden acceptance needs the browser to report a real
# window-visibility change through document.hidden, and the automation browser reports
# document.hidden as false for a window it simultaneously reports as minimized, with both
# visibility-suppressing launch arguments verifiably removed. That is a property of the
# automation browser rather than of this runner or of the product, so the observation
# overhead keeps running headed here while the hidden acceptance is recorded as not run.
# The visibility backoff itself stays covered by webui/src/state/sync.test.ts and the
# visibilitychange case in webui/tests/browser.spec.ts. See docs/webui-integration-matrix.md.
# Lock order: the cooperative CI flock, then the host `gpu-lock` that development sessions on
# this runner take (`/usr/local/bin/gpu-lock`, a flock on `/tmp/gpu-lock-$(id -u)/lock`; the
# runner is the same uid with PrivateTmp=no), then the fail-closed precheck below. Without
# the middle lock a development session's GPU job made the nvidia-smi check fail the gate
# (run 36982300183, #2111). On a host without gpu-lock the gate behaves as before. When the
# lock is not free within 600 seconds the step fails closed and prints its holder.
#
# The gate body is written to one script so it can run either under `gpu-lock run` or
# directly. Its first command, `activity_gate_host.py precheck` (#1949), stops a server an
# earlier run of this job leaked, or anything left working in a verifier run directory
# (orphaned, or its RUNNER_TEMP directory deleted), waits up to 600 seconds for the host to be quiet enough to measure (the
# instantaneous runnable-task count, not a load average: a concurrent compile on this host
# moved the score by more than the 2 percent budget), and then fails closed, naming each
# process, if any GPU compute process remains. The step timeout covers both 600-second lock
# waits and the 600-second quiet wait on top of the 45 minutes the gate itself was given.
- name: Run Activity performance gate with the hidden acceptance deferred
timeout-minutes: 75
env:
MLXCEL_REQUIRE_MODELS: '1'
run: |
set -euo pipefail
mkdir -p "$HOME/.cache/mlxcel/ci-locks"
gate_script="$RUNNER_TEMP/webui-activity-gate.sh"
export WEBUI_ACTIVITY_GATE_STARTED="$RUNNER_TEMP/webui-activity-gate.started"
rm -f "$WEBUI_ACTIVITY_GATE_STARTED"
cat > "$gate_script" <<'GATE'
set -euo pipefail
: > "$WEBUI_ACTIVITY_GATE_STARTED"
echo "Activity gate body started at $(date -u +%FT%TZ), $(( $(date +%s) - WEBUI_ACTIVITY_WAIT_START ))s after the host GPU lock wait began"
python3 scripts/webui/activity_gate_host.py precheck \
--installed-dir "$RUNNER_TEMP/mlxcel-webui-installed" \
--run-root "$RUNNER_TEMP" \
--host-state "$RUNNER_TEMP/webui-activity-host-state.json" \
--quiet-timeout 600
full_evidence="$RUNNER_TEMP/webui-activity-performance-full.json"
summary_evidence="$RUNNER_TEMP/webui-activity-performance-summary.json"
python3 scripts/webui/verify_activity_performance.py \
--server-bin "$RUNNER_TEMP/mlxcel-webui-installed/mlxcel-server-webui" \
--model "$RUNNER_TEMP/mlxcel-webui-activity-fixtures/qwen3-0.6b-4bit" \
--checkpoint-revision "ci-fixtures/qwen3-0.6b-4bit-v1" \
--source-sha "${{ github.sha }}" \
--features cuda,webui \
--evidence "$full_evidence" \
--work-dir "$RUNNER_TEMP" \
--repo-root "$GITHUB_WORKSPACE" \
--node-bin "$(command -v node)" \
--virtual-display \
--activity-timeout 1800 \
--pid-file "$RUNNER_TEMP/webui-activity-gate.pid" \
--host-state "$RUNNER_TEMP/webui-activity-host-state.json" \
--perf-mode visible-only-headed
python3 scripts/webui/summarize_activity_evidence.py \
--input "$full_evidence" \
--output "$summary_evidence" \
--expect-hidden-native not-run
GATE
(
# This flock serializes cooperating CI steps only; gpu-lock below covers development sessions on the host, and the precheck at the top of the gate body fails closed on any other GPU compute work.
flock -w 600 -x 9 || { echo "::error::timed out waiting for cooperative WebUI Activity GPU lock"; exit 1; }
WEBUI_ACTIVITY_WAIT_START="$(date +%s)"
export WEBUI_ACTIVITY_WAIT_START
if command -v gpu-lock >/dev/null 2>&1; then
echo "Host gpu-lock holder at $(date -u +%FT%TZ): $(gpu-lock status); waiting up to 600s for it"
rc=0
gpu-lock run --tag webui-activity --wait 600 -- bash "$gate_script" || rc=$?
if [ "$rc" -ne 0 ] && [ ! -e "$WEBUI_ACTIVITY_GATE_STARTED" ]; then
echo "::error::could not obtain the host gpu-lock within 600s (gpu-lock exit $rc); holder: $(gpu-lock status)"
exit 1
fi
exit "$rc"
fi
echo "gpu-lock is not installed on this runner; running the Activity gate without the host GPU lock"
bash "$gate_script"
) 9>"$HOME/.cache/mlxcel/ci-locks/webui-activity-gpu.lock"
# The verifier's server, display, node and browser start their own sessions, so neither a
# step timeout, a cancellation nor a killed verifier reaches them, and a survivor holds GPU
# memory into the next job (#1949). Reap by the pid file the verifier writes and by this
# job's identity. Anything reaped is a warning, a survivor fails this job instead of the
# next one, and nothing to do is the normal outcome.
- name: Reap Activity gate processes
if: always() && hashFiles('scripts/webui/activity_gate_host.py') != ''
timeout-minutes: 5
run: |
python3 scripts/webui/activity_gate_host.py reap \
--installed-dir "$RUNNER_TEMP/mlxcel-webui-installed" \
--run-root "$RUNNER_TEMP" \
--pid-file "$RUNNER_TEMP/webui-activity-gate.pid"
- name: Upload installed-artifact WebUI evidence
if: always()
uses: actions/upload-artifact@v7
with:
name: webui-installed-artifact-${{ runner.os }}
path: ${{ runner.temp }}/webui-installed-evidence.json
if-no-files-found: error
- name: Upload reverse-proxy WebUI evidence
if: always()
uses: actions/upload-artifact@v7
with:
name: webui-reverse-proxy-${{ runner.os }}
path: ${{ runner.temp }}/webui-reverse-proxy-evidence.json
if-no-files-found: warn
- name: Upload Activity performance summary evidence
if: always()
uses: actions/upload-artifact@v7
with:
name: webui-activity-performance-${{ runner.os }}
path: ${{ runner.temp }}/webui-activity-performance-summary.json
if-no-files-found: warn
# A fail-closed gate whose diagnosis stays on the runner cannot be acted on from a PR.
# The harness writes its browser log and a full evidence JSON with the failing condition;
# publish both whenever the gate runs, not only when it succeeds.
- name: Upload Activity performance diagnostics
if: always()
uses: actions/upload-artifact@v7
with:
name: webui-activity-performance-diagnostics-${{ runner.os }}
path: |
${{ runner.temp }}/webui-activity-performance-full.json
${{ runner.temp }}/webui-activity-host-state.json
${{ runner.temp }}/mlxcel-activity-performance-*/activity-performance.log
${{ runner.temp }}/mlxcel-activity-performance-*/*.log
if-no-files-found: warn
mlx-pin:
name: MLX pin extraction
needs: changes
if: needs.changes.outputs.mlx_pin == 'true'
runs-on: ubuntu-latest
permissions:
contents: read
steps:
- uses: actions/checkout@v7
with:
persist-credentials: false
# The pinned MLX commit has one home, GIT_TAG in
# src/lib/mlx-cpp/CMakeLists.txt, and two independent parsers read it:
# mlxcel-core's build.rs in Rust (build_support/mlx_pin.rs, unit-tested by
# mlxcel-mlx-pin) and this script in awk, for release.yml's "Validate MLX
# build cache" steps. The shell half accepts strictly less than the Rust
# half: it needs GIT_TAG as the first token of its own line, after the
# GIT_REPOSITORY line, with no closing parenthesis in between. A CMake
# reformat that leaves every local build and `make verify` green can
# therefore break every release build, and release.yml is the only other
# place this script runs, so without this job the breakage would first
# appear when a release is cut. Issue #1047.
#
# `mlxcel-mlx-pin` compiles in under a second (a leaf crate with no
# production role, only `std` plus a `tempfile` dev-dependency), so it
# runs first here rather than only in nightly-verify.yml's `make verify`.
# Its `tests/cross_parser.rs` shells out to mlx_pinned_commit.sh itself
# and asserts that whenever the shell parser accepts an input, the Rust
# parser resolves to the identical commit for that same input: the two
# unit-test files below each cover one parser in isolation and cannot
# catch a same-input value mismatch on their own.
- uses: dtolnay/rust-toolchain@1.97.1
- name: Run the Rust pin parser's unit and cross-parser tests
run: cargo test -p mlxcel-mlx-pin
- name: Extract the pinned MLX commit
run: scripts/ci/mlx_pinned_commit.sh
# The other half of the contract: the shell parser must also *reject*
# everything the Rust parser rejects. A value it accepts and the Rust
# parser refuses is not harmless, because release.yml purges every MLX
# build cache that does not match what this script prints and the build
# then fails on the same file the purge was justified by.
- name: Check the pinned MLX commit parser rejects malformed input
run: scripts/ci/mlx_pinned_commit_test.sh
cross-repo-refs:
name: cross-repo refs
if: github.event_name == 'pull_request'
runs-on: ubuntu-latest
permissions:
contents: read
steps:
- uses: actions/checkout@v7
with:
persist-credentials: false
fetch-depth: 0
# Advisory: lists bare 3+-digit '#NNN' added on this PR so unqualified
# upstream refs (-> org/repo#NNN) and any leaked private-repo numbers are
# caught in review. Does not fail the build (a genuine lablup/mlxcel #N is
# fine); run locally with STRICT=1 to gate. See CONTRIBUTING.md.
- name: Flag unqualified cross-repository references
# Pass the base ref via env (not interpolated into the shell) so a branch
# name can never reach shell parsing; the script reads BASE_REF as input.
env:
BASE_REF: origin/${{ github.base_ref }}
GITHUB_TOKEN: ${{ github.event.pull_request.head.repo.full_name == github.repository && github.token || '' }}
run: |
bash scripts/ci/check_cross_repo_refs_test.sh
python3 scripts/ci/check_cross_repo_refs.py
# ============================================================================
# OpenXLA feature compile (self-hosted GB10)
# ============================================================================
# Nothing else in CI compiles any XLA feature. `deny`, `fmt` and the other
# jobs here never build the crate, `pipeline-parallel-ci.yml` runs clippy on
# default features, and `nightly-verify.yml` uses `metal,accelerate`. Two
# defects reached `main` through that gap: a non-exhaustive `match` in the
# OpenXLA serve worker after a new `ModelRequest` variant landed, and a dead
# re-export that `-D warnings` would have rejected. Both were found by hand
# while rebasing an unrelated branch.
#
# Why this runs on the self-hosted GB10 runner rather than a hosted one:
# `xla-iree` links a C shim against a prebuilt IREE runtime, which
# `scripts/iree/setup-cuda.sh` provisions by building the runtime from source
# against a pinned revision. That is far too slow to do per PR from scratch,
# but the GB10 runner already holds the build under
# `~/.cache/mlxcel/iree-cuda-<version>` and the script is idempotent, so a
# warm runner reuses it and only a fresh one pays the one-time cost.
# `xla-diagnostics` additionally implies `cuda`, which points at the same
# runner.
#
# Deliberately NOT covered, so the gap is recorded rather than assumed shut:
# - Linking. `cargo check` never invokes the linker, so a regression in the
# IREE link recipe in `build.rs` (see #1274/#1275) is invisible here. The
# `xla-link` job immediately below closes that gap for the CUDA/IREE
# recipe; see its comment for what it covers and what it still does not.
# - `cargo test`. This is a compile gate; running the XLA suites needs a
# GPU that is not contended with development work on the same host.
# - Clippy-only lints. This job runs `cargo check`, and `RUSTFLAGS: "-D
# warnings"` denies rustc lints only, so `dead_code` and `unused_imports`
# are now gated but `collapsible_if`, `needless_match`,
# `too_many_arguments`, `while_let_loop` and the rest of clippy's own
# lints are not gated by anything in CI under these feature sets. The
# `clippy` job above covers default features only. That half of the
# backlog cleared by #1304 can therefore regrow silently; closing it means
# running clippy here, which was left out of #1304 deliberately.
# - The path dependencies' own test targets. There is no `default-members`
# (see the note at the top of `Cargo.toml`), so `cargo check` run here
# resolves to `-p mlxcel` and `--all-targets` expands within that package
# only. `mlxcel-core`, `mlxcel-surgery` and `mlxcel-xla` are compiled as
# plain library dependencies, so their `#[cfg(test)]` modules are never
# built. That leaves `mlxcel-xla`'s unit tests ungated by every job,
# because the `clippy` job above builds default features and the crate is
# default-off there. `cargo clippy -p mlxcel-xla --all-targets` covers
# them, but only by hand.
# - macOS and `IREE_DIST` builds of the same features. Neither has a runner
# with the required distribution.
xla-compile:
name: OpenXLA feature compile
needs: changes
if: >-
github.repository == 'lablup/mlxcel'
&& (needs.changes.outputs.rust == 'true' || contains(github.event.pull_request.labels.*.name, 'ci:full'))
runs-on: GB10
permissions:
contents: read
timeout-minutes: 120
env:
# The value shipped for this runner's architecture, so the gate compiles
# what ships rather than whatever the runner's GPU implies. Since issue
# #1943 auto-detection produces this same `121` on a GB10; the pin remains
# because a gate should name the list it compiles instead of inferring it.
# The hardware NVFP4/MXFP4 converters come from the per-source injection in
# `src/lib/mlx-cpp/CMakeLists.txt`, not from this list.
MLX_CUDA_ARCHITECTURES: "121"
# Every rustc lint, not just the `unused_imports` this job started with.
# The narrower value existed because the XLA feature combination carried a
# dead-code backlog that would have made the job red on arrival; #1304
# cleared it. See the note above for what `-D warnings` on `cargo check`
# still does not cover.
#
# Blast radius, deliberately: `RUSTFLAGS` reaches every unit cargo builds
# from a path, so the `mlxcel-core`, `mlxcel-surgery` and `mlxcel-xla`
# library units and the build scripts are gated here too, and a warning in
# any of them reds a job named for OpenXLA. Registry and git dependencies
# cannot, because cargo compiles those with `--cap-lints allow`. The
# toolchain is pinned exactly in `rust-toolchain.toml`, so a new rustc
# release cannot turn this red on its own; bumping that pin is what has to
# re-clear these two feature sets.
RUSTFLAGS: "-D warnings"
steps:
- uses: actions/checkout@v7
with:
persist-credentials: false
- name: Use a persistent target directory
run: |
# Separate from the release job's directory so a CI check cannot
# invalidate a release build's cache, or be slowed by it.
CARGO_TARGET="$HOME/.cargo-target/mlxcel-xla-ci"
mkdir -p "$CARGO_TARGET"
echo "CARGO_TARGET_DIR=$CARGO_TARGET" >> $GITHUB_ENV
- name: Provision the IREE runtime
run: |
# Idempotent: reuses `~/.cache/mlxcel/iree-cuda-<version>` when the
# runtime is already built, and builds it once when it is not.
bash scripts/iree/setup-cuda.sh
# The script emits shell `export VAR=value` lines; GITHUB_ENV wants
# bare `VAR=value`, and without stripping the prefix the variables
# would be named "export IREE_CUDA_HOME" and build.rs would abort
# claiming no IREE distribution is configured.
bash scripts/iree/setup-cuda.sh --env | sed 's/^export //' >> "$GITHUB_ENV"
- name: Compile the production XLA feature set
run: cargo check --features cuda,xla-iree --all-targets
- name: Compile the diagnostics XLA feature set
run: cargo check --no-default-features --features xla-diagnostics --all-targets
# ============================================================================
# OpenXLA feature link (self-hosted GB10)
# ============================================================================
# `xla-compile` above never invokes the linker, so it cannot catch a
# regression in the IREE link recipe in `build.rs`. That is not hypothetical:
# #1274 was exactly a link failure, and `cargo check --features
# cuda,xla-iree --all-targets` passed cleanly on the same tree that could not
# link a single integration test. It reached `main` and was found by hand
# while validating an unrelated PR. The fix, #1275, appends a second `-lc`
# after the IREE archives in `build.rs`, because rustc emits its own `-lc`
# before anything `cargo:rustc-link-arg` can append. This job is a direct
# reproducer of that failure class.
#
# What it links, and why: `cargo test --release --features cuda,xla-iree
# --test xla_prepared_prefill --no-run`. `build.rs` emits the IREE recipe
# through `cargo:rustc-link-arg`, which applies to every linked artifact of
# the crate, so any linked target exercises the same recipe `mlxcel-server`
# does, and an integration test is the direct reproducer because #1274 was an
# integration-test link failure. Cargo also builds the package's `[[bin]]`
# targets whenever an integration test is selected, so `mlxcel-server` and
# the other three binaries are linked here as well; the costs measured below
# already include them, which makes `cargo build --release --features
# cuda,xla-iree --bin mlxcel-server` a strict subset of this command rather
# than a cheaper alternative to it. `--no-run` links without executing, so
# this job never touches the GPU and never contends with development work on
# the shared GB10 host.
#
# Profile: `--release` is mandatory, not a choice. On this host the debug
# profile cannot link these targets at all: it fails with hundreds of
# `relocation truncated to fit: R_AARCH64_CALL26` errors against ordinary
# `libstd`/`compiler_builtins` symbols, because the unoptimized binary
# exceeds the AArch64 direct-branch range.
#
# Trigger: path-filtered via the `changes` job's `xla_link` output, not every
# Rust PR and not scheduled. `build.rs` holds the recipe, `scripts/iree/**`
# pins the IREE distribution whose archive set the recipe names, and
# `src/lib/mlxcel-xla/build.rs` with its `csrc/**` sources builds the shim
# object whose undefined symbols those archives resolve, so those are the
# causal surface. `rust-toolchain.toml` is on the same causal path because
# the #1274 failure was entirely about where rustc places its own `-lc`
# relative to appended `rustc-link-arg` entries, which a toolchain bump can
# move. `.github/workflows/ci.yml` is included so edits to
# this job itself are testable. A schedule was rejected because it decouples
# the failure from the PR that caused it, which is the exact failure mode of
# #1274 (found by hand later, not by CI). Running on every Rust PR was
# rejected because the release link is minutes of shared-runner time and the
# overwhelming majority of Rust PRs cannot affect the link line.
#
# `src/lib/mlxcel-core/build.rs` and `src/lib/mlx-cpp/CMakeLists.txt` are on
# the causal path as well. The former lists the CUDA libraries every
# `--features cuda` link resolves `libmlx.a` against; the latter holds the MLX
# pin, and a pin bump can add one without touching any build script. That is
# how ml-explore/mlx#4208's cuSOLVER dependency arrived with the 81ba1c6a bump
# (#1769): MLX links it PRIVATE, `cargo check` never links, and no job on the
# PR linked CUDA. A pin bump invalidates `mlxcel-core`'s MLX build, so those
# runs pay the cold figure below rather than the warm one.
#
# Measured cost of this job's link command on the GB10 runner: 15m40s from a
# purged `CARGO_TARGET_DIR`, most of which is MLX's CUDA sources compiling
# from scratch rather than the link itself, and 6m to 7m30s warm. Warm is the
# figure that matters, because the trigger below fires on `build.rs`, which
# invalidates the `mlxcel` crate but not `mlxcel-core`'s MLX build. For
# comparison, `xla-compile`'s `cargo check` over the same feature set is
# about a minute warm, which is why the two are separate jobs on separate
# triggers.
#
# `RUSTFLAGS` is deliberately unset here, unlike `xla-compile`. This job
# gates the link, not lints; `xla-compile` owns the lint policy. Leaving it
# unset means a red run here is unambiguously a link failure, not a lint
# failure wearing a link job's name.
#
# What the gate was demonstrated on: dropping `-l:libflatcc_parsing.a` from
# the `IREE_CUDA_HOME` recipe in `build.rs`. On that tree `cargo check
# --features cuda,xla-iree --all-targets` still passed in 59 seconds while
# this job's link command failed with undefined references to
# `flatcc_verify_*`, which is the same shape as #1274 and the exact gap this
# job exists to close.
#
# #1274's own defect, removing the second `-lc`, was tried as the control
# first and did not fail the link on the tree this job was validated against.
# Why it did not is unresolved, so do not read that run as evidence the entry
# is dead: against the pinned runtime, `nm` still reports `__stack_chk_guard`
# undefined in 176 objects of `libiree_runtime_unified.a`, including the
# `call.c.o` named in #1274; the symbol is still undefined in `libc.so.6`;
# and it is still defined only in `ld-linux-aarch64.so.1`. Every precondition
# the `build.rs` comment records holds today, so leave that entry alone until
# a control that actually reproduces says otherwise.
#
# Deliberately NOT covered, so the gap is recorded rather than assumed shut:
# - The `IREE_DIST` recipe (`build.rs:180-220`) and the macOS
# `IREE_MACOS_HOME` recipe (`build.rs:38-90`). Both remain unverified by
# any link, on any machine, because no runner in this repository holds
# either distribution. Only the `IREE_CUDA_HOME` recipe this job exercises
# is covered.
# - Anything outside the path filter above. Because the trigger is
# path-filtered rather than universal, a link regression arriving through
# a path the filter does not name (a new dependency pulling in a
# conflicting native library, for instance) is still not caught at PR
# time.
# - GPU execution. `--no-run` links only; the XLA test suites still have no
# CI coverage, per the `xla-compile` comment above.
# - A run where cargo finds nothing to redo. The target directory persists
# across PRs, so a green run does not by itself prove a link happened:
# this job's own first run on the PR that added it finished in 12 seconds
# with `Finished release profile in 0.14s`. That is correct, because
# cargo relinks exactly when the fingerprint moves, and `build.rs`
# declares `rerun-if-env-changed` for all three IREE variables so a
# version bump does move it. The narrow hole is a `scripts/iree/**` edit
# that changes how the runtime is built without moving `IREE_VERSION`:
# the script is idempotent, so nothing rebuilds and nothing relinks.
xla-link:
name: OpenXLA feature link
needs: changes
if: >-
github.repository == 'lablup/mlxcel'
&& (needs.changes.outputs.xla_link == 'true' || contains(github.event.pull_request.labels.*.name, 'ci:full'))
runs-on: GB10
permissions:
contents: read
# The other jobs on this runner are seconds warm; this one is minutes, so
# it is the first here that can hold GB10 long enough to matter. Without a
# group, three pushes to a `build.rs` PR queue three link jobs of up to
# 120 minutes each on the one runner that also serves clippy for every Rust
# PR and the release build. Keyed per PR, so PRs never cancel each other.
#
# `main` is exempt for the same reason the workflow-level group exempts it:
# a push to `main` carries no pull request number, so the key falls back to
# `refs/heads/main` and every merge shares one group. An unconditional flag
# would then let a second merge cancel this job mid-link on the branch a
# release builds from, which is the one run that has to finish.
concurrency:
group: xla-link-${{ github.event.pull_request.number || github.ref }}
cancel-in-progress: ${{ github.ref != 'refs/heads/main' }}
timeout-minutes: 120
env:
# The value shipped for this runner's architecture, so the gate links what
# ships rather than whatever the runner's GPU implies. Since issue #1943
# auto-detection produces this same `121` on a GB10; the pin remains
# because a gate should name the list it compiles instead of inferring it.
# The hardware NVFP4/MXFP4 converters come from the per-source injection in
# `src/lib/mlx-cpp/CMakeLists.txt`, not from this list.
MLX_CUDA_ARCHITECTURES: "121"
steps:
- uses: actions/checkout@v7
with:
persist-credentials: false
- name: Use a persistent target directory
run: |
# Its own directory, not shared with `xla-compile` or the release
# job: a link job emits release artifacts and build-script output
# that a check-only cache does not, and sharing would have the two
# jobs repeatedly invalidate each other.
CARGO_TARGET="$HOME/.cargo-target/mlxcel-xla-link-ci"
mkdir -p "$CARGO_TARGET"
echo "CARGO_TARGET_DIR=$CARGO_TARGET" >> $GITHUB_ENV
- name: Provision the IREE runtime
run: |
# Idempotent: reuses `~/.cache/mlxcel/iree-cuda-<version>` when the
# runtime is already built, and builds it once when it is not.
bash scripts/iree/setup-cuda.sh
# The script emits shell `export VAR=value` lines; GITHUB_ENV wants
# bare `VAR=value`, and without stripping the prefix the variables
# would be named "export IREE_CUDA_HOME" and build.rs would abort
# claiming no IREE distribution is configured.
bash scripts/iree/setup-cuda.sh --env | sed 's/^export //' >> "$GITHUB_ENV"
- name: Link the OpenXLA integration test
run: cargo test --release --features cuda,xla-iree --test xla_prepared_prefill --no-run
# ============================================================================
# CUDA sm_70 compile (self-hosted GB10)
# ============================================================================
# Every other CUDA job in this repository compiles for sm_121, and
# release.yml ships `80;86;89;90a;100;120` on x86_64 and `90a;100;121` on
# aarch64. Nothing below sm_80 is
# compiled anywhere, so an arch-conditional break on Volta reaches `main`
# unnoticed and is found by whoever next builds from source on one. That is
# not hypothetical for this repository: epic #1536 adds `cc < 8` fallbacks to
# the quantized-matmul overlays precisely because the sm_80+ paths are wrong
# there, and every one of those branches is compiled only when some job asks
# nvcc for a pre-Ampere target.
#
# Compile-only, and deliberately so. The realistic failure mode is a
# compile-time one: a `__CUDA_ARCH__` guard that does not cover cc 7, a
# CUTLASS type that has no pre-Ampere instantiation, a `bf16` intrinsic with
# no sm_70 implementation. Catching those needs nvcc pointed at sm_70, not a
# Volta card, and no GPU runner is involved here. The runtime half stays
# uncovered, and the decision to leave it uncovered is recorded in
# `docs/benchmark_results/volta-sm70-baseline-2026-08-31.md` (issue #1538)
# together with what a GPU-backed Volta job would and would not add: there is
# no Volta runner, no release artifact targets sm_70, and #1537 already turns
# an architecture mismatch into a named startup error instead of an opaque
# CUDA load failure.
#
# Path-filtered on `cuda_arch`, which is narrower than `rust`: a cold MLX
# build for a new architecture is measured in tens of minutes, and only the
# MLX overlays, the pinned commit and the build scripts can move what nvcc
# compiles. Its own persistent target directory, because the architecture
# list is part of the build-script fingerprint and sharing a directory with
# an sm_121 job would make the two rebuild MLX against each other on every
# run.
cuda-sm70-compile:
name: CUDA sm_70 compile
needs: changes
if: >-
github.repository == 'lablup/mlxcel'
&& (needs.changes.outputs.cuda_arch == 'true' || contains(github.event.pull_request.labels.*.name, 'ci:full'))
runs-on: GB10
permissions:
contents: read
# Same reasoning as `xla-link`: minutes warm, tens of minutes cold, on the
# one runner that also serves clippy for every Rust PR and the release
# build. Keyed per PR, so PRs never cancel each other.
#
# `main` is exempt, as in `xla-link` and at the workflow level: pushes to
# `main` carry no pull request number, so they share the `refs/heads/main`
# key, and an unconditional flag would let one merge cancel the previous
# merge's compile before it proves the shipped architecture still builds.
concurrency:
group: cuda-sm70-${{ github.event.pull_request.number || github.ref }}
cancel-in-progress: ${{ github.ref != 'refs/heads/main' }}
timeout-minutes: 120
env:
# Volta. Plain `70`, not `70a`: architecture-specific suffixes only exist
# from sm_90 on, and build.rs forwards this string to CMake verbatim.
MLX_CUDA_ARCHITECTURES: "70"
steps:
- uses: actions/checkout@v7
with:
persist-credentials: false
- name: Use a persistent target directory
run: |
# Not shared with any sm_121 job: the architecture list reaches CMake
# through the build script, so two jobs at different architectures
# sharing one directory would invalidate each other's MLX build every
# run and pay a cold nvcc pass each time.
CARGO_TARGET="$HOME/.cargo-target/mlxcel-cuda-sm70-ci"
mkdir -p "$CARGO_TARGET"
echo "CARGO_TARGET_DIR=$CARGO_TARGET" >> $GITHUB_ENV
- name: Check whether the toolkit can target sm_70
id: toolkit
run: |
# CUDA 13 removed Volta: nvcc there fails with
# `nvcc fatal : Unsupported gpu architecture 'compute_70'` before it
# compiles anything. The CUDA runners currently carry 13.x, so this
# gate cannot run on them, and a job that simply failed would red-light
# every PR touching these paths for a toolkit limitation rather than a
# code defect. Probe instead, and skip loudly. The moment a CUDA 12.x
# runner joins the fleet this job starts gating for real with no edit.
set -euo pipefail
# Same resolution order as the cuobjdump lookup below and as
# build.rs: CUDA_HOME, then the conventional location, then PATH.
# nvcc is not on PATH on this runner either.
NVCC="${CUDA_HOME:-/usr/local/cuda}/bin/nvcc"
if [ ! -x "$NVCC" ]; then
NVCC=$(command -v nvcc || true)
fi
if [ ! -x "${NVCC:-}" ]; then
echo "nvcc not found at ${CUDA_HOME:-/usr/local/cuda}/bin/nvcc or on PATH" >&2
exit 1
fi
VERSION=$("$NVCC" --version | sed -n 's/.*release \([0-9.]*\).*/\1/p' | head -1)
if "$NVCC" --list-gpu-arch 2>/dev/null | grep -qx compute_70; then
echo "supported=true" >> "$GITHUB_OUTPUT"
echo "CUDA $VERSION targets sm_70; gating for real."
else
echo "supported=false" >> "$GITHUB_OUTPUT"
echo "CUDA $VERSION cannot target sm_70; skipping the sm_70 build." | tee -a "$GITHUB_STEP_SUMMARY"
{
echo ""
echo "**sm_70 gate skipped.** This runner's CUDA toolkit is $VERSION, and CUDA 13 removed Volta (sm_70) support, so nvcc rejects \`compute_70\` outright. Nothing was compiled for Volta and nothing was verified. This is a toolkit limitation, not a passing gate. Building mlxcel for Volta requires a CUDA 12.x toolchain; see the CUDA architecture selection section of docs/installation.md."
} >> "$GITHUB_STEP_SUMMARY"
fi
- name: Compile the CUDA feature set for sm_70
if: steps.toolkit.outputs.supported == 'true'
run: cargo check --features cuda --all-targets
- name: Verify the emitted cubins are sm_70
if: steps.toolkit.outputs.supported == 'true'
run: |
# `cargo check` still runs the build script, so MLX is compiled and
# archived even though nothing links. Asserting on the archive turns
# "the environment variable was set" into "nvcc actually emitted
# pre-Ampere code", which is what the job claims to gate. It also
# catches the silent trap described in #1538: build.rs auto-detects
# the host and falls back to `90a` when it cannot, so a job whose
# architecture list failed to reach CMake would otherwise pass here
# having compiled nothing for Volta.
set -euo pipefail
# cuobjdump ships with the CUDA toolkit, next to nvcc, and neither is
# on PATH on the GB10 runner: a PATH-first lookup is what failed on
# the first run of the block-float job. Resolve the toolkit the way
# `src/lib/mlxcel-core/build.rs` resolves nvcc, CUDA_HOME first and
# then the conventional location, with PATH as a last resort rather
# than as the primary, so this reads the same toolkit that compiled
# the archive. Not found stays a hard failure: a skip here would make
# the gate pass while checking nothing.
CUDA_BIN="${CUDA_HOME:-/usr/local/cuda}/bin"
CUOBJDUMP="$CUDA_BIN/cuobjdump"
if [ ! -x "$CUOBJDUMP" ]; then
CUOBJDUMP=$(command -v cuobjdump || true)
fi
if [ ! -x "${CUOBJDUMP:-}" ]; then
echo "cuobjdump not found at $CUDA_BIN/cuobjdump or on PATH; it ships with the CUDA toolkit next to nvcc" >&2
exit 1
fi
echo "cuobjdump: $CUOBJDUMP"
LIB=$(find "$CARGO_TARGET_DIR/debug/build" -path '*/out/build/lib/libmlx.a' | head -1)
if [ -z "$LIB" ]; then
echo "libmlx.a not found under $CARGO_TARGET_DIR/debug/build" >&2
exit 1
fi
echo "archive: $LIB"
ARCHES=$("$CUOBJDUMP" --list-elf "$LIB" | grep -oE 'sm_[0-9]+a?' | sort -u)
echo "architectures present: $(echo "$ARCHES" | tr '\n' ' ')"
if [ "$ARCHES" != "sm_70" ]; then
echo "expected sm_70 only, found: $ARCHES" >&2
exit 1
fi
# ============================================================================
# CUDA block-float converter (self-hosted GB10)
# ============================================================================
# Compiles the shipped Blackwell list and then proves two different things
# about it: that MLX's hardware NVFP4/MXFP4 converter really reached the
# archive, and that the MXFP4/NVFP4 quantization tests still pass with it
# compiled in. Before issue #1934 neither held anywhere. Every list this
# repository built named Blackwell plainly, the converters in
# mlx/backend/cuda/quantized/nvfp4_quantize.cuh are gated on
# `__CUDA_ARCH_SPECIFIC__`, and so the tests that were supposed to cover
# block-float quantization exercised only the scalar CUTLASS fallback.
#
# It also proves the converter reached *only* there. The architecture lists
# stay plain and `src/lib/mlx-cpp/CMakeLists.txt` gives the extra image to
# `fp_quantize.cu` alone, so an architecture-specific image anywhere else
# means the injection was replaced by a target-level list and every kernel on
# every Blackwell host now runs different machine code. `cuda-arch-lists`
# above catches that spelling in the YAML; this catches it in the artifact,
# which also covers a list that never reaches CMake at all (build.rs
# auto-detects and falls back to `90a` when it cannot see the host). That
# trap is not hypothetical; it is the same one `cuda-sm70-compile` asserts
# against, for the same reason.
#
# Gate on SASS, never on PTX. The injection emits a cubin and no PTX, and the
# PTX that is emitted comes from the plain `compute_121` pass, which takes the
# fallback arm, so `cuobjdump --dump-ptx | grep cvt.rn.satfinite.e2m1x2`
# returns zero on a *correct* build.
#
# Path-filtered on `cuda_arch`, deliberately and with a known gap: a PR that
# changes only the Rust quantization code or its tests does not run this job,
# because a cold MLX build is measured in tens of minutes on the one
# self-hosted runner that also serves clippy for every Rust PR. What the filter does cover is everything that can change *which*
# device code is compiled, which is what the gate is about. Its own
# persistent target directory, because the architecture list is part of the
# build-script fingerprint and sharing with a differently-pinned job would
# make the two rebuild MLX against each other on every run.
cuda-blockfloat:
name: CUDA block-float converter
needs: changes
if: >-
github.repository == 'lablup/mlxcel'
&& (needs.changes.outputs.cuda_arch == 'true' || contains(github.event.pull_request.labels.*.name, 'ci:full'))
runs-on: GB10
permissions:
contents: read
# Same reasoning as `xla-link` and `cuda-sm70-compile`: minutes warm, tens
# of minutes cold, on the one runner that also serves clippy for every Rust
# PR and the release build. Keyed per PR, so PRs never cancel each other,
# and `main` is exempt so a merge cannot cancel the run that proves the
# shipped architecture still carries the converter.
concurrency:
group: cuda-blockfloat-${{ github.event.pull_request.number || github.ref }}
cancel-in-progress: ${{ github.ref != 'refs/heads/main' }}
timeout-minutes: 180
env:
# The list releases build for this runner's architecture. Keep it
# identical to the aarch64 release job's Blackwell half: a gate that
# compiles something else proves something else.
MLX_CUDA_ARCHITECTURES: "121"
steps:
- uses: actions/checkout@v7
with:
persist-credentials: false
- name: Use a persistent target directory
run: |
CARGO_TARGET="$HOME/.cargo-target/mlxcel-cuda-blockfloat-ci"
mkdir -p "$CARGO_TARGET"
echo "CARGO_TARGET_DIR=$CARGO_TARGET" >> $GITHUB_ENV
- name: Build the block-float quantization tests
# `test-fast` inherits release but drops fat LTO, which is what the
# other CUDA test builds in this repository use: these tests drive MLX
# numerics, so an unoptimized build is not merely slower but
# impractically so.
run: cargo test -p mlxcel-core --profile test-fast --features cuda --lib --no-run
- name: Verify the hardware fp4 converter is compiled, and only where intended
run: |
set -euo pipefail
# cuobjdump ships with the CUDA toolkit, next to nvcc, and neither is on PATH
# on the GB10 runner: a PATH-first lookup is what failed on this job's first
# run. Resolve the toolkit the way `src/lib/mlxcel-core/build.rs` resolves
# nvcc, CUDA_HOME first and then the conventional location, with PATH as a
# last resort rather than as the primary, so this reads the same toolkit that
# compiled the archive. Not found stays a hard failure: a skip here would make
# the gate pass while checking nothing.
CUDA_BIN="${CUDA_HOME:-/usr/local/cuda}/bin"
CUOBJDUMP="$CUDA_BIN/cuobjdump"
if [ ! -x "$CUOBJDUMP" ]; then
CUOBJDUMP=$(command -v cuobjdump || true)
fi
if [ ! -x "${CUOBJDUMP:-}" ]; then
echo "cuobjdump not found at $CUDA_BIN/cuobjdump or on PATH; it ships with the CUDA toolkit next to nvcc" >&2
exit 1
fi
echo "cuobjdump: $CUOBJDUMP"
LIB=$(find "$CARGO_TARGET_DIR/test-fast/build" -path '*/out/build/lib/libmlx.a' | head -1)
if [ -z "$LIB" ]; then
echo "libmlx.a not found under $CARGO_TARGET_DIR/test-fast/build" >&2
exit 1
fi
echo "archive: $LIB ($(stat -c %s "$LIB") bytes)"
# One checkout can hold several build directories, one per feature set and
# architecture list, and `head -1` above picks whichever comes first. Assert
# that the one being read was configured with the list this job pins, so the
# gate cannot pass by measuring a neighbour.
BUILD_DIR="$(dirname "$LIB")/.."
CONFIGURED=$(sed -n 's/^MLX_CUDA_ARCHITECTURES:[^=]*=//p' "$BUILD_DIR/CMakeCache.txt")
echo "configured with: $CONFIGURED"
if [ "$CONFIGURED" != "$MLX_CUDA_ARCHITECTURES" ]; then
echo "archive was configured with [$CONFIGURED], not the job's [$MLX_CUDA_ARCHITECTURES]" >&2
exit 1
fi
WORK="${RUNNER_TEMP:-/tmp}/blockfloat-$$"
mkdir -p "$WORK"
( cd "$WORK" && ar x "$LIB" fp_quantize.cu.o qmv.cu.o fp_qmv.cu.o )
for obj in fp_quantize.cu.o qmv.cu.o fp_qmv.cu.o; do
if [ ! -f "$WORK/$obj" ]; then
echo "$obj not found in $LIB; MLX's quantized backend did not build as expected" >&2
exit 1
fi
done
# 1. The converter is compiled. Zero means the architecture-specific image
# never reached nvcc and MLX compiled the scalar fallback, which is the
# state issue #1934 found in every artifact this project shipped.
#
# Scoped to the one translation unit that can carry the instruction rather
# than to the whole archive: `--dump-sass` on the 178 MB archive takes about
# ten minutes and this takes under a second, for the same verdict.
"$CUOBJDUMP" --dump-sass "$WORK/fp_quantize.cu.o" > "$WORK/sass.txt"
FP4=$(grep -c 'F2FP.SATFINITE.E2M1' "$WORK/sass.txt" || true)
echo "F2FP.SATFINITE.E2M1 occurrences in fp_quantize.cu.o: $FP4"
if [ "$FP4" -eq 0 ]; then
echo "no hardware fp4 conversion instruction in libmlx.a: MLX compiled the scalar fallback" >&2
exit 1
fi
# The fp8 converter is not gated on an architecture-specific target, so it is
# present either way. Printing it beside the fp4 count makes "nothing compiled
# at all" read differently from "the fp4 gate is off".
echo "F2FP.SATFINITE.E4M3 occurrences (ungated control): $(grep -c 'F2FP.SATFINITE.E4M3' "$WORK/sass.txt" || true)"
# 2. The converter is compiled ONLY there. This is the half that separates the
# per-source injection from naming Blackwell `121a` in the architecture
# list. The list form gives every object an architecture-specific image, and
# on a matching device that is the image the driver loads, so every kernel
# on every Blackwell host would run different machine code than it does
# today. Decode is supposed to be untouched by a converter fix.
for obj in qmv.cu.o fp_qmv.cu.o; do
IMAGES=$("$CUOBJDUMP" --list-elf "$WORK/$obj" | grep -oE 'sm_[0-9]+[af]?' | sort -u | tr '\n' ' ')
echo "$obj images: $IMAGES"
if printf '%s' "$IMAGES" | grep -qE 'sm_[0-9]+[af]'; then
echo "$obj carries an architecture-specific image; decode was not supposed to change (issue #1934)" >&2
exit 1
fi
done
ARCH_SPECIFIC=$("$CUOBJDUMP" --list-elf "$LIB" | grep -cE 'sm_[0-9]+a' || true)
echo "architecture-specific cubin images in the whole archive: $ARCH_SPECIFIC"
if [ "$ARCH_SPECIFIC" -ne 1 ]; then
echo "expected exactly one architecture-specific image in libmlx.a, the one in fp_quantize.cu.o, found $ARCH_SPECIFIC" >&2
exit 1
fi
# 3. Forward compatibility is untouched. Plain `compute_121` PTX is what a
# later Blackwell revision JITs from, and the injection emits a cubin only,
# so this count must not move.
PTX=$("$CUOBJDUMP" --dump-ptx "$LIB" | grep -cE '^\.target sm_121$' || true)
echo "forward-JIT-capable plain PTX images (.target sm_121): $PTX"
if [ "$PTX" -eq 0 ]; then
echo "no plain compute_121 PTX in libmlx.a; the architecture list is not what this job pinned" >&2
exit 1
fi
# 4. The assumption the per-source injection rests on: `nvfp4_quantize.cuh` is
# the only place in MLX that gates code on an architecture-specific target.
# If a pin bump adds another, a quantized GEMM say, the injection silently
# stops covering the arch-specific surface, and this is the only place that
# would notice.
MLX_SRC="$BUILD_DIR/_deps/mlx-src/mlx"
if [ ! -d "$MLX_SRC" ]; then
echo "MLX source tree not found at $MLX_SRC; cannot check the arch-specific surface" >&2
exit 1
fi
GATED=$(grep -rl '__CUDA_ARCH_SPECIFIC__\|__CUDA_ARCH_FAMILY_SPECIFIC__' "$MLX_SRC" | sed "s|$MLX_SRC/||" | sort | tr '\n' ' ' | sed 's/ $//')
echo "MLX files gated on an architecture-specific target: $GATED"
if [ "$GATED" != "backend/cuda/quantized/nvfp4_quantize.cuh" ]; then
echo "the arch-specific surface of the pinned MLX has changed; the per-source injection in src/lib/mlx-cpp/CMakeLists.txt covers only fp_quantize.cu, so decide whether to extend it or move to a target-level architecture list (issue #1934)" >&2
exit 1
fi
rm -rf "$WORK"
- name: Run the MXFP4 and NVFP4 quantization tests
# `global_scale` selects the six `ffi_tests` cases that drive
# `quantize_weights_with_mode` with `mxfp4` and native NVFP4 sidecars,
# which is every test in this crate that reaches MLX's block-float
# quantization kernels. `shapeless_compile_audit_harness` touches them
# too but carries `#[ignore]` for a hardware audit script, so naming it
# here would add a skipped line and no coverage.
#
# `--test-threads=1` is mandatory for CUDA on this runner: parallel MLX
# tests abort regardless of the change under test.
run: |
cargo test -p mlxcel-core --profile test-fast --features cuda --lib -- \
--test-threads=1 \
global_scale
# ============================================================================
# ROCm build, link and generate (self-hosted gfx1151)
# ============================================================================
# Nothing else in CI compiles the ROCm feature at all. `cargo check` jobs do
# not carry it, the CUDA jobs point at NVIDIA runners, and the ROCm half of
# every MLX pin bump now lives in `src/lib/mlx-cpp/patches-rocm/`. Without
# this job a bump that breaks the HIP overlay reaches `main` unnoticed, which
# is exactly how the CUDA link regression in #1274 got in: `cargo check` was
# green on a tree that could not link a single test.
#
# So this job builds AND links AND generates. Each step catches a class the
# one before it cannot:
# - `cargo build --release --features rocm` compiles the overlay and links
# the binary, which `cargo check` would not do.
# - The generate step runs a real forward pass on the GPU, which catches a
# kernel that compiles and links but launches nothing (the `searchsorted`
# and `hadamard` gaps in #1825 were of that shape).
# - The device assertion catches a silent fall back to the CPU, which would
# otherwise make a broken GPU path look like a passing job. It reads the
# `Runtime device:` line and not the kernel-backend one: those two are
# different questions, and a run forced onto the CPU still reports
# `custom kernel backend: rocm`, because that names the backend MLX
# resolved rather than the device the run used (#1805).
#
# The runner is a shared host with other GPU tenants, so the job takes one
# small checkpoint, one short generation and a persistent target directory
# rather than a clean build per run.
rocm-build:
name: ROCm build and generate
needs: changes
# `vars.ROCM_CI_ENABLED` is an activation switch, not a feature flag. A job
# targeting a self-hosted runner that does not exist is not skipped by
# GitHub, it is queued, for up to 24 hours. While it sits there the whole
# workflow run stays in progress, which means no job in that run has
# readable logs, and the check reads as "still running" to a reviewer. So
# every PR touching the HIP overlay would pay for the runner being absent.
# Parking the job behind a repository variable makes it genuinely inert
# until someone registers the runner and sets the variable to `true`, and
# `rocm-runner-status` below says which state it is in.
if: >-
github.repository == 'lablup/mlxcel'
&& (needs.changes.outputs.rocm == 'true' || contains(github.event.pull_request.labels.*.name, 'ci:full'))
&& vars.ROCM_CI_ENABLED == 'true'
runs-on: [self-hosted, rocm-gfx1151]
permissions:
contents: read
# Generous because a cold runner builds MLX from source; a warm one reuses
# the persistent target directory and finishes in minutes.
timeout-minutes: 180
env:
# Set explicitly rather than detected. `rocminfo` needs `/dev/kfd`, which
# a containerised runner may not see, and detection failing would turn a
# build error into a confusing "no GPU agent" abort. This is also the
# value the host actually is, so a mismatch between the two is a real
# finding rather than a CI artefact (#1805).
MLX_ROCM_ARCHITECTURES: gfx1151
steps:
- uses: actions/checkout@v7
with:
persist-credentials: false
- name: Use a persistent target directory
run: |
# Separate from any other job's directory, so a ROCm check cannot
# invalidate another build's cache or be slowed by it.
CARGO_TARGET="$HOME/.cargo-target/mlxcel-rocm-ci"
mkdir -p "$CARGO_TARGET"
echo "CARGO_TARGET_DIR=$CARGO_TARGET" >> "$GITHUB_ENV"
- name: Report the ROCm toolchain
run: |
# Recorded in the log so a failure can be attributed to a host change
# rather than to the PR. Not a gate: a missing `rocminfo` is fine
# because the build does not need it when the architecture is set.
hipconfig --version || true
rocminfo 2>/dev/null | grep -E 'Name:|gfx' | head -20 || true
- name: Build and link with the ROCm feature
run: cargo build --release --features rocm
- name: Cache CI fixture weights
id: ci-fixture-cache
uses: actions/cache@v6
with:
path: models/qwen3-0.6b-4bit
key: ${{ runner.os }}-ci-fixture-qwen3-0.6b-4bit-v1
- name: Download CI fixture
if: steps.ci-fixture-cache.outputs.cache-hit != 'true'
env:
GH_TOKEN: ${{ secrets.GITHUB_TOKEN }}
CI_FIXTURE_TAG: ci-fixtures/qwen3-0.6b-4bit-v1
CI_FIXTURE_ASSET: qwen3-0.6b-4bit.tar.gz
run: |
mkdir -p models
tmp_dir="$(mktemp -d)"
gh release download "$CI_FIXTURE_TAG" \
--pattern "$CI_FIXTURE_ASSET" \
--dir "$tmp_dir"
tar -xzf "$tmp_dir/$CI_FIXTURE_ASSET" -C models/
rm -rf "$tmp_dir"
- name: Generate on the GPU and assert the device and the output
env:
# Same checkpoint the fixture step above unpacks.
MLXCEL_ROCM_SMOKE_MODEL: models/qwen3-0.6b-4bit
run: |
# The body lives in `scripts/ci/rocm_smoke.sh` so this job and
# `make verify-rocm-smoke` cannot drift. That matters more than usual
# here: until a gfx1151 runner exists this job never runs, so the
# local target is the only thing exercising the script, and a copy
# kept in this file would rot unnoticed (lablup/mlxcel#1811).
bash scripts/ci/rocm_smoke.sh
# ============================================================================
# ROCm CI status (hosted)
# ============================================================================
# `rocm-build` above is parked until a `rocm-gfx1151` runner exists and
# `vars.ROCM_CI_ENABLED` is `true`. Parking it silently would be its own
# trap: a reviewer would see no ROCm check at all on a PR that changes the
# HIP overlay and reasonably conclude one ran. This job always reports which
# state the ROCm gate is in.
#
# It is advisory and never fails a PR. Whether a runner is attached is an
# infrastructure fact, not a defect in the change under review.
rocm-ci-status:
name: ROCm CI status
needs: changes
if: github.repository == 'lablup/mlxcel' && needs.changes.outputs.rocm == 'true'
runs-on: ubuntu-latest
permissions:
contents: read
timeout-minutes: 5
steps:
- name: Report the state of the ROCm gate
env:
ROCM_CI_ENABLED: ${{ vars.ROCM_CI_ENABLED }}
run: |
# Reports the variable and nothing else. An earlier version also
# queried runner registration, which needs an `administration` scope
# that GITHUB_TOKEN has no such key for: adding it is not a weaker
# permission, it is an invalid workflow file, and GitHub then refuses
# to run any job in the file at all. The variable is what gates the
# job, so it is the honest thing to report.
if [ "${ROCM_CI_ENABLED:-}" = "true" ]; then
echo "ROCm CI is enabled; the 'ROCm build and generate' job in this run is the authority."
exit 0
fi
echo "::warning title=ROCm gate parked::This PR changes paths that affect the ROCm build, but 'ROCm build and generate' did not run: vars.ROCM_CI_ENABLED is not 'true'. Register a runner labelled rocm-gfx1151 and set that repository variable to 'true' to activate it (lablup/mlxcel#1811)."
echo "ROCm gate parked: vars.ROCM_CI_ENABLED is not 'true'."
echo "To activate: register a self-hosted runner labelled rocm-gfx1151, then set the repository variable ROCM_CI_ENABLED to 'true'."
# ============================================================================
# Workflow lint (hosted)
# ============================================================================
# Every path filter used to list `.github/workflows/ci.yml`, so a workflow
# edit re-ran the heavy self-hosted jobs. The stated reason was to validate a
# CI change before merge, but it was aimed at the wrong failure.
#
# Workflow changes break in two ways. A schema or expression error makes
# GitHub reject the whole file and run nothing, so re-running heavy jobs
# catches nothing: the run just reports "This run likely failed because of a
# workflow file issue." That is not hypothetical, it happened on #1989 from a
# single invalid `permissions` key. And an error in a job's own script, which
# only running that job can catch, and which the `ci:full` label now opts
# into deliberately.
#
# This job covers the first class properly, in seconds, on a hosted runner.
# `actionlint` rejects the exact key that broke #1989 by name. The repository
# is clean against it, which took a `.github/actionlint.yaml` listing the
# self-hosted runner labels: without that, seven `runner-label` findings would
# have made this job red on arrival and therefore worthless.
workflow-lint:
name: workflow lint
needs: changes
if: github.repository == 'lablup/mlxcel' && needs.changes.outputs.workflows == 'true'
runs-on: ubuntu-latest
permissions:
contents: read
timeout-minutes: 10
steps:
- uses: actions/checkout@v7
with:
persist-credentials: false
- name: Run actionlint
env:
# Pinned. A floating version can add a check and turn this red on a
# PR that changed nothing about workflows, which is how a lint job
# stops being trusted.
ACTIONLINT_VERSION: 1.7.7
run: |
curl -sSLo actionlint.tar.gz \
"https://github.com/rhysd/actionlint/releases/download/v${ACTIONLINT_VERSION}/actionlint_${ACTIONLINT_VERSION}_linux_amd64.tar.gz"
tar xzf actionlint.tar.gz actionlint
./actionlint -version
# Severity floor, not a suppression. The hosted image ships a shell
# linter, so actionlint also checks every `run:` block, and this
# repository carries 21 info and 4 warning findings there that predate
# the job (mostly SC2086 quoting and SC2016 single-quoted
# expressions). Landing red on all of them would make the job
# something to ignore or delete. Errors are what break a script, and
# at this floor the job is still load-bearing: the only two
# error-level findings when it was written were its own, from a
# comment line whose first word was read as a linter directive.
# Clearing the info and warning backlog is separate work.
#
# Do not start a comment line in any `run:` block with that linter's
# name; it is parsed as a directive and reports SC1072/SC1073.
./actionlint -shellcheck='shellcheck --severity=error'
# ============================================================================
# Self-hosted gate advisory (hosted)
# ============================================================================
# The jobs on the shared runners no longer re-run when a workflow file
# changes. That is the point, but left silent it would be its own trap: a PR
# that rewrites the script inside `cuda-sm70-compile` would show no CUDA check
# at all, and a reviewer could reasonably read that as "it passed".
#
# So say it. Advisory only, never failing, because failing would put the cost
# back: a documentation PR that happens to touch a workflow comment would be
# blocked until someone ran the heavy jobs it cannot affect. Whoever does want
# them run adds the `ci:full` label, which every self-hosted job now honours.
self-hosted-gate-advisory:
name: self-hosted gate advisory
needs: changes
if: >-
github.repository == 'lablup/mlxcel'
&& github.event_name == 'pull_request'
&& needs.changes.outputs.workflows == 'true'
&& !contains(github.event.pull_request.labels.*.name, 'ci:full')
runs-on: ubuntu-latest
permissions:
contents: read
timeout-minutes: 5
steps:
- name: Note that the self-hosted jobs were not re-run for this workflow change
run: |
echo "::warning title=Self-hosted jobs not re-run::This PR changes .github/workflows/**, which no longer re-runs the jobs on the shared GB10 and gfx1151 runners (lablup/mlxcel#1992). They run when the paths they actually gate change. If this PR edits the script inside one of those jobs, add the 'ci:full' label to run them and validate it before merge."
echo "Self-hosted jobs were gated on the paths they cover, not on this workflow change."
echo "Add the 'ci:full' label to force them."