You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
[Feature] nodeconfig: match nodes by label selector (nodelabelselector), not only by exact name #3068
Allow a nodeconfig entry in the device-plugin config.json to select nodes by a Kubernetes label selector instead of (or in addition to) an exact node name:
Proposed matching order in readFromConfigFile(): (1) exact name match, (2) the first entry whose nodelabelselector matches the node's labels (list order = precedence, log a warning if several match), (3) the "*" entry, (4) global defaults. Existing configs keep working unchanged.
Why is this needed:
Today readFromConfigFile() matches os.Getenv(NODE_NAME) == val.Name only, so per-node settings (operatingmode, devicesplitcount, devicememoryscaling, migstrategy, filterdevices) need one entry per node name. This breaks down whenever node names are not stable:
Autoscaled or reprovisioned nodes (Karpenter, EKS managed node groups, AWS Capacity Block rollovers) get new names, so every rollover means editing values, helm upgrade, and restarting the plugin.
"name": "*" is widely used as a fallback but is not documented, and it is not a wildcard.
This was requested in #2040 and closed in the 2026-08 backlog cleanup without an implementation. We hit the same problem operating a GPU-as-a-Service platform whose node groups run in different operatingmodes and are replaced regularly, and currently run a small out-of-tree controller that renders the nodeconfig list from a node label and restarts the plugin pod when an entry changes. It works, but the plugin can do this natively at start-up by reading its own Node object (it already has a Kubernetes client), which removes the extra component and the race with helm upgrade rewriting the ConfigMap.
Scope note: this changes only which entry applies to a node. The per-node fields themselves are already node-scoped and are reported to the scheduler through the hami.io/node-nvidia-register annotation, so no scheduler change is needed (unlike the per-node memoryFactor discussion in #2295).
Anything else we need to know?:
Implementation sketch: add NodeLabelSelector *metav1.LabelSelector (json:"nodelabelselector") to DevicePluginConfigs.Nodeconfig, fetch the node's labels once at start-up, match with metav1.LabelSelectorAsSelector, document the order above in charts/hami/README.md and docs/develop/dynamic-mig.md, and add unit tests for name match, selector match, multiple matches, and the "*" fallback. No RBAC change: the device-plugin ClusterRole already has get on nodes (it patches the register annotation).
Reacting to label changes after start-up (watch + self-restart) can be a follow-up; a DaemonSet restart covers it for now.
What would you like to be added:
Allow a
nodeconfigentry in the device-pluginconfig.jsonto select nodes by a Kubernetes label selector instead of (or in addition to) an exact node name:{ "nodeconfig": [ { "name": "gpu-node-01", "operatingmode": "hami-core", "devicesplitcount": 10 }, { "nodelabelselector": { "matchLabels": { "gpu.example.com/pool": "mig" } }, "operatingmode": "mig", "devicesplitcount": 10, "migstrategy": "none" }, { "nodelabelselector": { "matchExpressions": [ { "key": "karpenter.k8s.aws/instance-gpu-name", "operator": "In", "values": ["t4", "l4"] } ] }, "operatingmode": "hami-core", "devicesplitcount": 4 }, { "name": "*", "operatingmode": "hami-core", "devicesplitcount": 10 } ] }Proposed matching order in
readFromConfigFile(): (1) exactnamematch, (2) the first entry whosenodelabelselectormatches the node's labels (list order = precedence, log a warning if several match), (3) the"*"entry, (4) global defaults. Existing configs keep working unchanged.Why is this needed:
Today
readFromConfigFile()matchesos.Getenv(NODE_NAME) == val.Nameonly, so per-node settings (operatingmode,devicesplitcount,devicememoryscaling,migstrategy,filterdevices) need one entry per node name. This breaks down whenever node names are not stable:helm upgrade, and restarting the plugin.hami-coremode and another inmigmode (the question in In a k8s cluster, there are both t4 cards and a100 cards. How to configure the t4 to use hami core and the a100 to use mig #1126) have to list each node by hand."name": "*"is widely used as a fallback but is not documented, and it is not a wildcard.This was requested in #2040 and closed in the 2026-08 backlog cleanup without an implementation. We hit the same problem operating a GPU-as-a-Service platform whose node groups run in different
operatingmodes and are replaced regularly, and currently run a small out-of-tree controller that renders thenodeconfiglist from a node label and restarts the plugin pod when an entry changes. It works, but the plugin can do this natively at start-up by reading its own Node object (it already has a Kubernetes client), which removes the extra component and the race withhelm upgraderewriting the ConfigMap.Scope note: this changes only which entry applies to a node. The per-node fields themselves are already node-scoped and are reported to the scheduler through the
hami.io/node-nvidia-registerannotation, so no scheduler change is needed (unlike the per-nodememoryFactordiscussion in #2295).Anything else we need to know?:
NodeLabelSelector *metav1.LabelSelector(json:"nodelabelselector") toDevicePluginConfigs.Nodeconfig, fetch the node's labels once at start-up, match withmetav1.LabelSelectorAsSelector, document the order above incharts/hami/README.mdanddocs/develop/dynamic-mig.md, and add unit tests for name match, selector match, multiple matches, and the"*"fallback. No RBAC change: the device-plugin ClusterRole already hasgetonnodes(it patches the register annotation).