Skip to content

Queue-aware GPU sharing capacity and RDNA concurrent-inference expectations for AMD #3144

Description

@moezdil

Summary

AMD GPU sharing on RDNA needs the scheduler to account for compute-queue (HQD) capacity, and the docs to set realistic expectations for concurrent inference. Today capacity is modeled from cores and memory only, but a real limit on gfx12 is the small number of user compute HQD slots and CP pipes.

Why

On gfx12 there are 4 user compute HQD slots by default (8 with amdgpu num_kcq=0) over 2 CP pipes. Each ROCm process uses several compute queues (Ollama about 5), so a few processes oversubscribe the slots and short synchronised decode kernels serialise across processes. This collapses concurrent inference throughput regardless of the CU mask or the VRAM slice.

Verified on real hardware (RX 9060 XT, gfx1200, ROCm 10.0)

Ollama 0.12.5-rocm, 256 tokens, two half-CU slices generating at once:

  • qwen2.5:3b: 71.9 tok/s alone in a slice, 12.3 tok/s each when both run (about 83% loss).
  • smollm2:1.7b: 105 alone, about 103 each concurrently (no meaningful loss).
  • num_kcq=0 (8 HQD slots) did not recover the qwen collapse on this firmware.

So concurrent-inference behavior is model and firmware dependent, and a slice alone is unaffected.

Proposal

  1. Let the scheduler read per-GPU compute-HQD capacity (surfaced by amd-device-plugin) and factor it into how many processes may share a card.
  2. Document RDNA expectations: two slices can work for light models but may collapse for heavier decode; the loss is driver/firmware dispatch contention, not the slicer.
  3. Raise per-process queue placement (same CP pipe) with AMD as an upstream limit.

Part of the AMD vGPU parity umbrella #3139. Node-side companion issue: Project-HAMi/amd-device-plugin (compute-queue capacity accounting).

Additions

  • Cap the number of sharing processes per GPU, not only the queues. Four slices collapsed to 11-17 tok/s each even with GPU_MAX_HW_QUEUES=1 sized to fit exactly 8 HQD slots, because gfx12 has only 2 CP pipes; more than about 2 processes per GPU is effectively a hard limit.
  • Do not co-locate a whole-GPU workload with sliced pods on one card: the whole-GPU pod has no dmem VRAM cap and can take the slices' memory (see Evaluate dmem cgroup controller as a fail-closed VRAM backend amd-hami-core#6).

Status

No activity

Activity on this issue will appear here.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions