Skip to content

[Feature] nodeconfig: match nodes by label selector (nodelabelselector), not only by exact name #3068

Description

@usr-bin-ksh

What would you like to be added:

Allow a nodeconfig entry in the device-plugin config.json to select nodes by a Kubernetes label selector instead of (or in addition to) an exact node name:

{
  "nodeconfig": [
    { "name": "gpu-node-01", "operatingmode": "hami-core", "devicesplitcount": 10 },
    { "nodelabelselector": { "matchLabels": { "gpu.example.com/pool": "mig" } },
      "operatingmode": "mig", "devicesplitcount": 10, "migstrategy": "none" },
    { "nodelabelselector": { "matchExpressions": [
        { "key": "karpenter.k8s.aws/instance-gpu-name", "operator": "In", "values": ["t4", "l4"] } ] },
      "operatingmode": "hami-core", "devicesplitcount": 4 },
    { "name": "*", "operatingmode": "hami-core", "devicesplitcount": 10 }
  ]
}

Proposed matching order in readFromConfigFile(): (1) exact name match, (2) the first entry whose nodelabelselector matches the node's labels (list order = precedence, log a warning if several match), (3) the "*" entry, (4) global defaults. Existing configs keep working unchanged.

Why is this needed:

Today readFromConfigFile() matches os.Getenv(NODE_NAME) == val.Name only, so per-node settings (operatingmode, devicesplitcount, devicememoryscaling, migstrategy, filterdevices) need one entry per node name. This breaks down whenever node names are not stable:

This was requested in #2040 and closed in the 2026-08 backlog cleanup without an implementation. We hit the same problem operating a GPU-as-a-Service platform whose node groups run in different operatingmodes and are replaced regularly, and currently run a small out-of-tree controller that renders the nodeconfig list from a node label and restarts the plugin pod when an entry changes. It works, but the plugin can do this natively at start-up by reading its own Node object (it already has a Kubernetes client), which removes the extra component and the race with helm upgrade rewriting the ConfigMap.

Scope note: this changes only which entry applies to a node. The per-node fields themselves are already node-scoped and are reported to the scheduler through the hami.io/node-nvidia-register annotation, so no scheduler change is needed (unlike the per-node memoryFactor discussion in #2295).

Anything else we need to know?:

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions