Umbrella for AMD device plugin work: RDNA WGP-aware CU slicing, real health, metrics, multi-GPU, and hook handling.
Done:
Real device health instead of always-healthy #28 per-GPU health, exporter health matched by BDF (feat(plugin): per-GPU health from the DRM device #87 )
Do not inject libamvgpu LD_AUDIT hook for whole-GPU requests #29 skip the libamvgpu LD_AUDIT hook for whole-GPU requests
besteffort allocator p2pWeights init fails on single-GPU nodes #30 besteffort allocator on single-GPU nodes
Skip CDNA partition sysfs reads on RDNA #31 skip CDNA partition sysfs reads on RDNA (verified on gfx1200)
Allocator policies beyond besteffort (spread/binpack) #35 spread and binpack allocator policies
Align and update Go module dependencies #43 align Go module dependencies (incl. kubelet v0.37.1)
Clean up code flagged by staticcheck #44 staticcheck cleanup
Review deadcode candidates for unreachable functions #53 dead code removal
RDNA WGP-aware CU allocation in AllocateN #21 -Tests for WGP-aware CU allocation #25 RDNA WGP-aware CU allocation (fix(cuallocation): allocate whole WGPs on RDNA #105 , feat(plugin): publish cuPerWGP so the scheduler accounts whole WGPs #107 ; verified with pods on gfx1201, scheduler side in fix(amd): round core requests up to whole WGPs HAMi#3174 )
Implement metricssvc GetGPUState and List #26 metricssvc exporter API
Validate ROCR_VISIBLE_DEVICES ordering on multi-GPU nodes #27 multi-GPU ROCR_VISIBLE_DEVICES and HSA_CU_MASK ordering (test(plugin): lock HSA_CU_MASK order to ROCR_VISIBLE_DEVICES on multi-GPU #100 ), APUs without a ROCr UUID (fix(plugin): allocate GPUs without a ROCr UUID by container-local index #101 ), allocator on GPUs without peer links (fix(allocator): weight GPU pairs without a direct link through the host #102 ), per-GPU memory limit (fix(plugin): set the memory limit per GPU on multi-GPU slices #103 ), per-GPU dmem cap (fix(plugin): let the dmem cgroup back the musl fail-closed check #106 , test in test(plugin): cover the per-GPU dmem cap #108 ); verified on an RX 9070 XT plus a Radeon iGPU in one node
Container Device Interface (CDI) support #34 CDI support (feat(plugin): inject GPUs through CDI #91 ; verified with pods on gfx1201)
Account for compute-queue (HQD) capacity when sharing AMD GPUs on gfx12 #54 compute-queue capacity: slots published (feat(plugin): publish compute queue slots in custominfo #88 , feat(plugin): make the GPU split count configurable #90 ), gfx12 defaults to 2 sharers per GPU (feat(plugin): default to 2 sharers per GPU on gfx12 #109 ); scheduler-side caps in Queue-aware GPU sharing capacity and RDNA concurrent-inference expectations for AMD HAMi#3144
CU mask is advisory: HSA_CU_MASK cannot be enforced from user space #55 CU mask is advisory: documented (docs: document that CU slices are cooperative #89 ); libamvgpu pins HSA_CU_MASK to the pod spec so setenv/os.environ cannot widen it (fix: pin HSA_CU_MASK to the pod spec value amd-hami-core#13 , shipped in chore: bump amd-hami-core to pin HSA_CU_MASK #111 ); a child exec'd without LD_AUDIT still escapes, which needs a CU ceiling in KFD
dmem VRAM cap on by default where the node supports it (feat(plugin): cap sliced VRAM with dmem by default where supported #110 )
A GPU whose DRM device can no longer be opened is marked unhealthy: verified by unbinding a passthrough iGPU from amdgpu (allocatable 20 -> 10, back to 20 after rebind)
In review:
Validation pending (needs other hardware or software):
/kind feature
Umbrella for AMD device plugin work: RDNA WGP-aware CU slicing, real health, metrics, multi-GPU, and hook handling.
Done:
HSA_CU_MASKto the pod spec so setenv/os.environ cannot widen it (fix: pin HSA_CU_MASK to the pod spec value amd-hami-core#13, shipped in chore: bump amd-hami-core to pin HSA_CU_MASK #111); a child exec'd without LD_AUDIT still escapes, which needs a CU ceiling in KFDIn review:
Validation pending (needs other hardware or software):
/kind feature