Skip to content

[Allocator 2/5 — after #71] Model GPU fabric, NUMA, MIG, and transfer contention - #73

Open
deehw wants to merge 2 commits into
codex/dynamic-resource-allocator-v2from
codex/allocator-advanced-topology
Open

[Allocator 2/5 — after #71] Model GPU fabric, NUMA, MIG, and transfer contention#73
deehw wants to merge 2 commits into
codex/dynamic-resource-allocator-v2from
codex/allocator-advanced-topology

Conversation

@deehw

@deehw deehw commented Sep 2, 2026

Copy link
Copy Markdown
Contributor

Stacked follow-up to #71. Adds schedulable GPU/MIG device and peer-link wire models; best-effort NVIDIA UUID, MIG, NUMA, PCIe, NVLink, and NIC discovery; joint GPU subset feasibility; new model-profile fabric/NUMA/MIG constraints; and cold-load ranking based on artifact size, transfer bandwidth, and active transfer contention. Advanced requirements fail closed on unknown topology. Includes 522 focused passing tests and operator docs. Base will be retargeted to main after #71 merges.

@deehw
deehw marked this pull request as draft September 3, 2026 00:57
@deehw deehw changed the title Allocator: model GPU fabric, NUMA, MIG, and transfer contention [Allocator 2/5 — after #71] Model GPU fabric, NUMA, MIG, and transfer contention Sep 3, 2026
@deehw
deehw marked this pull request as ready for review September 3, 2026 03:11
@deehw

deehw commented Sep 3, 2026

Copy link
Copy Markdown
Contributor Author

Qualification update: physically validated on C (2x RTX 4090) and D (1x RTX 4090). This found and fixed idle PCIe generation underestimation; C now reports the native PHB topology at 23.628 GB/s rather than 3.0 GB/s. Focused tests, lint, CI, and the 4,164-test combined suite all pass. Ready for review.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant