Grid's allocator decides which configured models should be resident on which computers, then converges the fleet toward that placement without taking a healthy last replica away. It is built for an AI intranet where capacity is heterogeneous and partly opportunistic: an always-on GPU server, an employee laptop, and an external inference endpoint do not have the same ownership or safety rules.
The allocator is experimental. It defaults to recommend mode and does not change a process until an operator explicitly selects automatic mode.
The global loop optimizes the grid. The local loop protects one computer. They exchange desired and actual state, but the local decision is authoritative when they disagree.
request counts, latency, queues host telemetry and local override
| |
v v
bounded demand forecast local protection loop
| |
v |
capacity-aware placement plan |
| |
v |
safe staged reconciliation <--------------------+
|
v
LOAD -> WARM -> READY -> DRAIN -> UNLOAD
^
|
heartbeat actual state + acknowledgements
The router and allocator have separate jobs. The router selects one ready engine for one live request. It never tells the allocator what to provision. The allocator independently observes the ordinary request lifecycle at the Grid boundary and changes future supply across the fleet.
For every completed or failed request, Grid derives bounded features in memory: endpoint family, requested and served model, text/image/video modality, approximate input and output units, one of a small configured workload vocabulary, service time, queue pressure, and outcome. Raw prompts, images, tool arguments, responses, API keys, and user identities are not retained. Observation is best effort and cannot delay or fail inference.
The controller keeps direct named-model pressure separate from workload demand. Named-model traffic scales that model normally. Unbound or automatic traffic contributes to a workload forecast such as coding, research, design, image, video, embedding, or general language work. At planning time, the allocator projects sustained workload demand onto configured models using each profile's workload scores, compatibility and resource footprint, cold-start time, and measured outcomes. This lets an inactive specialist receive a bounded canary without creating a loaded-only feedback loop. The model that happened to serve an automatic request contributes outcome evidence, but the router's fallback choice never becomes direct demand for that model.
If an iterative client supplies X-Grid-Affinity-Key, the same already-hashed session signal also
lets the allocator learn real cross-workload sequences. Five repeated transitions from one
workload to the same next workload are required, and the confidence calculation includes every
observed departure from the source. For example, repeated image→video campaign sessions can warm a
ComfyUI video model while the next image request is still running. Adjacent image and video traffic
from unrelated users is not sequence evidence and cannot move a model. The raw affinity key is
never retained; workflow cursors and transitions contain only a double-hashed key, are capped at
512 active workflows and 4,096 transitions, and expire after one hour. Out-of-order or simultaneous
completion callbacks update aggregate demand but cannot establish ordering. grid allocator status labels these proactive decisions with the learned workload edge and its confidence.
If a correlation-only prediction cannot yet receive a serving slot, the authoritative plan may
stage one policy-bounded cache-only advisory on a future-compatible host. This requires an exact
operator-authorized source, SHA-256, declared artifact size, authenticated free disk, and managed
load capability. Reconciliation downloads and verifies the artifact but never warms it, consumes
VRAM, or drains an incumbent; direct and baseline demand therefore keep all serving and eviction
authority. A later real-demand placement can reuse the cache and skip the network-transfer phase.
By default, speculative transfers must leave at least 10 GiB free; operators can tune or disable
that planner reserve. Missing free-disk telemetry never authorizes a predictive transfer. Demanded
placements retain their ordinary disk-admission policy, so this safety floor cannot block real
work. Counterfactual portfolio evaluations do not emit these operational advisories.
Portfolio admission normally requires three observations, preventing a cheap one-off request from churning a large model into memory. Device-time evidence can cross the gate earlier: at least 1.5 offered-concurrency units—arrival rate multiplied by measured service time, plus queued work—is already enough pressure to justify one candidate. A queued minute-long video job can therefore start its model immediately, while one short text request remains below the gate. Both thresholds are persisted and reported by allocator status.
Image and video work between 1.0 and 1.5 offered-concurrency units may earn an earlier canary-only placement, but only after a fleet-level spare-capacity proof. The controller first plans direct and ordinary-evidence workloads, treats unknown model-slot limits as having no hypothetical headroom, preserves slots for one unseen workload and one node failure, requires a second compatible host, and requires capacity beyond one feasible copy of every enabled catalog model. This distinguishes “the model fits on a host” from “the fleet can safely explore it.” Weak evidence is capped at one replica even when the raw concurrency calculation asks for more. A fresh successful response from that exact model/workload/artifact validates the canary and removes the cap; direct demand or the ordinary evidence threshold does the same. Saturated fleets retain the normal evidence gate, while genuinely overprovisioned fleets can react to a long sparse job one planning interval earlier.
Measured model/workload outcomes use confidence that reaches full weight after twenty fresh requests; separately labeled quality reaches full weight after eight fresh evaluations. Both decay with independent seven-day half-lives: a fresh latency-only request cannot revive stale quality. Evidence is keyed by model, workload, and immutable artifact SHA-256, so a replacement artifact starts uncertain and a late response from the previous revision cannot update the new one. Legacy shared-timestamp quality is discarded conservatively during restore. A six-point bounded optimism bonus lets an equally suitable, currently feasible cold candidate earn a canary, then falls to zero as fresh evidence accumulates. The bonus is smaller than meaningful configured-suitability differences and a preemption-only candidate pays a larger penalty, so uncertainty can spend spare capacity but cannot manufacture eviction authority. Status reports service and quality evidence age/freshness, effective sample count, confidence, quality confidence, and the remaining exploration bonus for every candidate.
Request latency is compacted into a bounded logarithmic histogram per time bucket. Aggregate SLO graduation therefore uses an approximate request-level p95: one slow outlier among ninety-nine fast requests remains tail evidence, but it no longer makes the entire minute appear slow.
Portfolio selection first scans the current fleet with the same hard runtime, backend, GPU, tag, data-tier, artifact, memory-headroom, model-slot, and colocation rules used by placement. An attractive model that no live node can host is excluded instead of suppressing a usable fallback. Among otherwise similar feasible candidates, startup time and current residency provide a bounded transition penalty without overwhelming measured quality or configured workload suitability. Status exposes every candidate's current-headroom feasibility, immutable compatibility, possible planner-authorized preemption path, eligible-node count, best host, startup estimate, transition penalty, and rejection reason. A candidate with an avoidable cold start must show a meaningful score improvement over a resident peer; after a justified switch the penalty reverses, providing state-dependent hysteresis without a stale controller-side lease. Portfolio canaries may use current headroom or request a planner-authorized replacement of stale speculative capacity. They cannot evict directly observed service or another speculative model's only active canary; configured/pinned baselines are likewise never relocation victims for speculative demand. Direct observed demand has broader relocation and preemption authority. Every path still requires the planner to prove victim priority, ownership, pins, active work, minimum residency, and capacity, and the reconciler drains before unloading. This bounded late binding lets a newly active workload replace an idle specialist instead of remaining permanently invisible behind a full model slot. When several occupied compatible hosts exist, the feasibility scan retains up to sixteen ranked preemption paths instead of treating the first host as the whole fleet. The controller selects the first path that preserves direct/baseline work and the active victim's required ready-replica floor. If a clean structural plan proves the recovered fleet can host one preferred model per active workload portfolio, one copy of each active model remains protected while an excess replica may be reclaimed to end whole-workload starvation. If that complete portfolio cannot fit—as on the scarce four-node fixture—the allocator normally retains the demand-derived floor. A mature beneficiary with at least 1.5× the incumbent's confidence-weighted device-time pressure may instead use one exact planner-proved replacement path. Pinned models, configured minimum replicas, and directly named demand remain absolute fences. The controller carries this authorization explicitly into the forecast, and an unvalidated replacement is capped at one canary. The portfolio evaluator counts the staged beneficiary as future service only when the authoritative plan names its preemption; intended capacity is never treated as physically free.
When two or more workload classes are active, Grid no longer picks each model independently. It starts from the evidence-backed choices and runs a deterministic bounded coordinate search over complete workload-to-model maps, evaluating each candidate portfolio with the authoritative fleet planner. Configured baselines and direct demand are preserved first. When the configured workload catalog is larger than the nominal fleet's model-slot count, Grid additionally searches bounded admitted model subsets. A workload may be explicitly deferred instead of every active workload creating a desired replica and leaving accidental model ordering to choose the losers. The scarce fleet objective maximizes distinct admitted workload coverage, then uses recent measured request failure to restore service among equally broad portfolios before comparing service-time-aware pressure and request coverage. This failure signal stays inside the bounded demand window and does not create durable service debt or bypass residency guards. The objective preserves spare slots when two portfolios serve the same demand, minimizes missing replicas, and then compares measured utility and transition cost. Resource pressure uses offered concurrency—the arrival rate multiplied by measured service time, plus queued work—so a long image or video job is not incorrectly treated as cheaper than a short embedding call merely because fewer jobs arrive. For an unbound workload, however, the queue has not yet been served by the counterfactual model. It can admit one policy-safe canary, but it is removed from that model's projected concurrency so one unroutable request cannot be counted both as workload backlog and as several speculative replicas. Repeated arrival/device-time evidence lifts the canary cap, while queue-driven scale-out comes from model-specific observations after real traffic reaches the selected model.
The subset search is ranked by cheap set coverage before its candidates are verified by the real planner, letting it cross multi-change local optima such as consolidating marketing/sales on an already-required general model and research/coding on one code model to free an embedding slot. Validated alternatives do not consume the uncertainty budget merely because their configured suitability is below a specialist. Nominally roomy fleets keep the pressure-first fast path, and a temporary outage does not switch portfolio policy. A shared generalist can therefore beat two slightly better specialists when only one model slot is available. Search considers at most four candidates per workload and 64 distinct portfolios, so catalog size cannot create an unbounded planning pass. Every counterfactual reuses one immutable workload-forecast snapshot and omits the executable plan identity digest; the single authoritative plan still hashes its complete inputs for generation fencing before any command can reach a node. An immediately repeated read-only status query reuses that portfolio bundle only when timestamp, node snapshots, model profiles, and the telemetry revision are identical. Status exposes the hit, and any new request or evaluation forces a fresh solve; reconciliation, admissions, commands, and lifecycle history remain live on every response. Fleet-wide topology option value is computed with host-to-model indexes and direct eligible-set intersections. This preserves the same constrained-host penalties without expanding every candidate/scarce-host/alternative-host triple on large 32–64 node planning passes. Counterfactual portfolio plans over the exact same timestamp, node objects, model profiles, and policy also share one bounded hard-topology context: compatibility, runtime footprint, future host sets, post-pin eligible-host counts, and fastest startup paths. Startup paths are additionally keyed by the complete validated learned warm/load timing maps. A new heartbeat/profile object, control timestamp, policy, or timing map invalidates the affected context; demand-only alternatives reuse it because they cannot alter those fleet facts.
Each bounded workload set reserves representation for its exploitation leader and the broadest
cross-workload candidate; a fifth-ranked generalist can therefore remain discoverable when four
narrow specialists would make every independently preferred portfolio infeasible.
Only one distinct model may differ from the exploitation-only portfolio because of uncertainty at a
time; this is an explicit fleet exploration budget, not one canary allowance per workload. A joint
portfolio may likewise introduce at most one preemption-only model per control tick, and receives no
coverage credit for it unless the evaluated full-fleet plan either places it or explicitly stages
its safe victim transition. This prevents individually plausible swaps from becoming an impossible
set of simultaneous promises on a saturated fleet. Status
shows the joint mapping, selected model set, and the model currently consuming that exploration
slot. It also reports one admission row per active unbound workload using the authoritative plan
and observed node state: ready, starting, planned, undersupplied, deferred,
capacity-contended, blocked-by-residency, infeasible, or awaiting-plan. Each row includes
the chosen model, offered concurrency, desired/planned/ready/missing replica counts, eligible hosts,
startup estimate, and the concrete reason. An intentionally omitted model is reported as deferred under current capacity; its demand and concurrency remain visible. A safe preemption candidate may
replace a sole speculative specialist when it continues every active workload that justified the
victim, while direct, pinned, and minimum replicas remain hard blockers. These rows explain the
existing decision; they do not run a second solver or alter router behavior. The controller snapshots
bounded demand and outcome state under the telemetry mutex, releases it, and performs planner-backed
search on that immutable snapshot; request completion telemetry is therefore not serialized behind
portfolio optimization.
Counterfactual portfolio demand is deliberately weaker than direct evidence: it may fill spare capacity but cannot evict a baseline or a directly demanded model. Workload-wide latency, queues, and errors remain workload evidence and are never copied into a candidate model as if that model had served the work. If direct demand later needs the hosts, the normal drain and unload guards reclaim the canary first. See ADR 0039 for the control-loop boundary.
The global loop consumes:
- model profiles: memory, optional immutable artifact SHA-256, authenticated artifact source and transfer-size ceiling, runtime/backend compatibility, replica bounds, priority, data tier, placement tags, failure-domain goal, co-location ceiling, pins, and cooldowns;
- host snapshots: usable memory, reserve, runtime/backend, lifecycle state, policy tags, cached and resident models, free/total artifact disk, concurrency, queue, measured throughput and latency, memory bandwidth, compute, heartbeat age, and actuator ownership;
- short-horizon demand: request rate, offered concurrency, queue depth, p95 latency, errors, and trend.
Demand history is bounded by time buckets and contains aggregate timings and counts, not prompts or responses. Bursts are folded into their bucket rather than truncated at a raw-request limit. A bucketed EWMA and positive trend term make the near-term forecast. The target replica count adds demand headroom and reacts to queue, latency, and error pressure while respecting per-model minimum and maximum bounds. Rising demand is also projected across the fastest eligible next replica's artifact-locality-aware load-plus-warm path, confidence-weighted and capped to a five-minute/2× horizon, so slow cold starts begin before the queue arrives without letting one noisy slope cause a fleet-wide load spike. Negative trends never accelerate scale-down. The model-demand tracker also learns mature groups of configured models that repeatedly become active in the same time buckets and directional model pairs that repeatedly activate one bucket apart. Current demand for one group member can prewarm a quiet peer. These model-level signals use a confidence-weighted historical rate ratio, require at least three supporting buckets and a 0.70 association/transition threshold, and exclude incomplete future buckets from transition failures. Cross-workload portfolio anticipation uses the stricter affinity-bound sequence proof described above; aggregate timing across unrelated users is never enough. Inferred demand is capped at twice the target's observed peak, takes the maximum rather than sum across at most 32 current sources, and never propagates transitively. Old target-local queue, latency, and error evidence is not refreshed by an association. Only configured, non-retiring model IDs create demand series, so the permissionless inference endpoint cannot grow controller state with arbitrary names.
Correlation-derived demand retains a separate observed-rate field. Within the same administrator priority class, required baseline replicas place first, then models with direct request/queue/SLO evidence, then inferred-only prewarms. This makes prediction opportunistic: it can use otherwise idle capacity, but an alphabetically earlier speculative model cannot evict the only slot available to real traffic. Older forecast producers without correlation lineage remain direct evidence for wire compatibility.
Placement is deterministic. Hard pins are reserved first and higher-priority classes fill before
lower-priority work. Within one priority class, constrained and larger models define each round's
order, but placement progresses one replica per model per round. Scarce capacity is therefore
max-min fair among equally important compatible models instead of being monopolized by the first
model ID before its peers receive a baseline. If capacity ends partway through one equal-share
round, the tie is resolved by attributable service harm: measured queue, latency/SLO ratio, errors,
concurrency, and observed request rate. Model IDs remain
only the final deterministic tie-break, so renaming two otherwise identical workloads cannot move
the scarce slot away from the workload with greater measured harm. If lower-priority managed residencies already occupy
all compatible capacity, Grid emits an explicit staged preemption: it first drains and unloads the
lowest-priority sufficient victim set, continues reporting the important model as
unsatisfied, and places it only after a later heartbeat proves that memory is actually free. The
same mechanism converges a host whose model ceiling was lowered below its live inventory. Each plan
stages at most 64 individual evictions by default; larger changes converge over later heartbeats
instead of producing an unbounded operational wave. Within a single-domain unpinned wave,
independent node-local victim sets are proven in one fleet scan and then consumed in disruption-cost
order; pins and multi-domain placement retain fresh searches. External,
manual, pinned, and minimum-residency-protected work is never bypassed. Correlation-only predictive
demand may use spare capacity but cannot trigger a destructive preemption; a configured baseline,
pin, or direct request/queue/SLO/error signal is required. Among equally low-priority choices, the
allocator first prefers failed, already-draining, and idle victims so urgent capacity does not wait
behind avoidable live work. It then prefers the set with the lowest learned warm-back cost, reducing
the delay of restoring displaced service after the burst. Required failure-domain diversity is reserved
before that disruption comparison, so several convenient victims in one rack cannot strand a critical model
that needs capacity across racks. A missing hard pin targets its exact node before either domain or
victim selection; freeing a different host that cannot satisfy the pin would be gratuitous disruption.
When a staged preemption names its beneficiary and that profile has an immutable operator-authorized
artifact source, reconciliation fetches and verifies the weights before draining the incumbent.
This disk-only CACHED state consumes no runtime memory or model slot, and WARM remains fenced until
a later authoritative plan assigns the released capacity; failed transfers retain ordinary bounded
backoff and never weaken drain safety.
Candidates otherwise prefer an existing ready residency, local cached weights, another failure
domain, measured throughput, and best-fit memory. Before measured
throughput exists, bounded memory-bandwidth and compute estimates break otherwise-cold ties; ready
and cached bonuses remain much larger at light demand, so hardware estimates do not cause
gratuitous migration. As offered concurrency exceeds one replica's target capacity, Grid increases
the value of measured performance and hardware speed by a bounded factor (at most 8×). Sustained
hot demand can therefore amortize a materially faster cold host, while a light workload still
prefers cached weights. This affects new placement only; it does not manufacture a migration after
the desired replica set is already healthy.
An uncached candidate must also fit its declared artifact-size ceiling in the node's authenticated
free-disk observation. The planner reserves those bytes cumulatively across every uncached desired
model on a multi-model host; several artifacts that fit independently cannot overcommit the same
free-space reading. Exact cached weights do not pay that disk cost again, and predictive transfers
start from the post-assignment disk ledger rather than racing demanded loads for space. The node
refreshes disk immediately before load, so a placement that became stale fails locally instead of
filling the filesystem; the next heartbeat lets the controller select another eligible host.
Placement also computes each desired model's future-compatible host set without pretending that
live memory is already free. A flexible small model is penalized on a host needed by a more
constrained model when it has another valid destination. A newly loaded incumbent remains sticky
for its configured minimum residency under this soft scarcity pressure, preventing adjacent control
ticks from reversing the same migration. A demanded model's sole feasible host is a hard override:
Grid may warm the incumbent's replacement immediately, while reconciliation still delays the drain
until minimum residency has elapsed. If an immediately free host is materially
worse after accounting for startup time, current work, hardware, and fit, Grid may
stage a proactive repack instead: it assigns the incumbent to its replacement host, proves that
replacement ready, drains and unloads the old copy, and only then warms the demanded model on the
released host. The target remains sticky while the victim is DRAINING; an intermediate heartbeat
cannot redirect the beneficiary to an avoidable cold or slower host. Marginal improvements use
immediate capacity, which bounds churn.
Ready, loading, and warming incumbents on full one-model hosts are indexed and ranked in one pass
when their failure domains are independent. Treating an in-progress incumbent as occupied prevents
the empty-host fast path from starting a duplicate cold load while its heartbeat is still
converging. Before opening new slots, the planner also seeds each target with healthy live
incumbents. If every target is already satisfied, that complete current placement is treated as a
feasibility witness: soft scarcity scoring cannot dismantle it and then return a less complete
greedy rebuild. A genuinely missing replica, hard policy change, or sole-host conflict still uses
ordinary relocation and bounded preemption. The same optimization applies to empty one-model hosts
only when one model remains in its priority class, preserving equal-priority sharing. Both cases
preserve the general scorer's exact result while avoiding a fleet-wide rescan for every replica on
large networks.
When several equal-priority models share an otherwise uniform empty fleet, Grid caches each static
candidate order but still consumes it one replica per model per allocation round; any shared-host or
domain interaction falls back to complete rescoring and bounded repacking.
Load, latency, host priority, cold-start time, and throttling lower a candidate's score. Managed
nodes report monotonic action duration in their authenticated acknowledgements. Successful warm
times are retained in bounded controller history and blended with the configured model estimate as
startup prior. Real missing-artifact fetches are marked separately from local cache verification;
their successful durations likewise become per-host load estimates, while verification-only loads
cannot teach the controller that a network transfer is nearly instant. A bounded eight-sample EWMA
becomes authoritative after four samples for that node/model/artifact revision. A checksum change
starts with the configured priors instead of inheriting optimistic timings from different weights.
Placement favors both faster cached hosts and faster cold-fetch hosts. Portfolio selection,
predictive prewarming, priority preemption, and mutation scheduling all use the fastest eligible
learned load-plus-warm path; a replica already warming does not pay the artifact-load phase twice.
An unknown host keeps the conservative configured fallback. Samples expire after 30 days so a
runtime, storage, or network upgrade can relearn. Invalid, non-finite, negative, or over-one-hour
reports are ignored rather than poisoning scheduling or receipt delivery. Persisted
failed warm/load attempts apply a bounded per-model penalty, allowing a healthy peer to be tried
after backoff instead of selecting the same broken cache forever; the failed node remains a fallback
when it is the only feasible target. Mutation acknowledgements and inventory snapshots are fenced
by observation time: a snapshot older than a successful action cannot trigger a duplicate drain or
unload, while a genuinely newer heartbeat that still reports the old state can start a new
lifecycle. Explicit pins, per-host model limits, compatibility policies,
and a feasible minimum failure-
domain count are hard constraints. A failure-domain shortfall or capacity shortage is reported
rather than hidden by overcommit. A throttled host exposes only its configured fraction of capacity
for new placement. Placement keeps two memory ledgers: current live processes consume incremental
make-before-break capacity including admission headroom, while the complete desired footprint must
fit the host's raw allocatable memory after reserve and thermal derating. An existing process may
therefore be re-admitted without requiring phantom free headroom, but reserve growth cannot leave an
unsafe collection of zero-incremental incumbents selected. The lowest-priority movable incumbent is
migrated first when a safe peer exists; without one, Grid reports the shortfall and preserves the
last replica instead of hiding either constraint. When greedy placement fragments capacity, a deterministic bounded backtracking
repair can evacuate and re-place several equal-priority replicas; unrelated or ineligible inventory
does not change that search budget or its result. On homogeneous empty-residency fleets—where moves
provably preserve resource use—aggregate free memory and model slots provide an admissible lower
bound, so an already saturated fleet fails fast. Runtime-specific memory and existing-residency
cases retain the full repair search because relocation can change their net footprint.
Request routing uses the same heterogeneous capacity evidence after placement. Among engines in the
same host-protection class that already serve the requested model, Grid compares active requests as
a fraction of each engine's effective concurrency limit rather than comparing raw request counts.
Its expected-completion estimate also includes advertised queued work, so an otherwise fast batched
vLLM engine with a backlog yields to a clear peer while active concurrency remains the hard admission
boundary.
This keeps a wide-batching vLLM server from appearing busier than a narrow llama.cpp engine merely
because it safely carries more simultaneous work. Missing capacity remains conservative raw load;
zero capacity remains closed. Throttled-host priority and hard admission limits still take
precedence over this load balance. Equivalent engines prefer the freshest lease, but a timestamp
inside the allowed future-skew window is clamped to zero age and cannot gain extra priority.
When private proxy measurements exist for the requested model, routing minimizes estimated
completion time: the incoming request's service wave is multiplied by a confidence- and
freshness-weighted latency EWMA. Weak or missing measurements blend toward comparable nodes'
median, measurements for other models are ignored, and expired evidence falls back completely.
For text generation, a bounded max_completion_tokens or max_tokens hint adds a model-throughput
lower bound to that estimate, allowing short requests to favor low latency and long generations to
favor high token throughput. Grid never inspects or stores prompt content for this decision.
Clients with multi-turn or iterative workloads may send an opaque X-Grid-Affinity-Key header.
Grid immediately hashes a printable key of at most 256 UTF-8 bytes and uses rendezvous hashing to
keep that model's requests on the same near-equivalent engine, preserving runtime KV/prompt caches
without a centralized session map. For automatic model requests, the allocator also double-hashes
this value and uses only repeated within-session workload transitions as proactive portfolio
evidence; the router still makes no provisioning decision. The raw key is neither retained nor
forwarded upstream. Host
protection and admission remain hard gates, and affinity considers only the best protection class
and routes whose estimated completion time is within 20% of the best available route. Adding or
removing an otherwise equivalent engine therefore remaps only the sessions assigned to the changed
engine; load, throttling, or failure can still move a session immediately.
Lease health is not the only routing signal. Grid keeps a private per-engine, per-model circuit breaker for outcomes observed by the proxy. A 429 opens a one-second cooldown immediately; two consecutive transport or 5xx failures open it, with exponential backoff capped at 30 seconds. An expired circuit admits a half-open probe and any 2xx response resets the streak. Caller-caused 4xx responses do not poison route health, and a broken model route on a multi-model vLLM server does not hide its healthy models. Circuit state never changes discovery inventory, never grants lifecycle authority, and is not persisted across a Grid-server restart. Grid also does not automatically replay a failed POST: the breaker redirects only subsequent requests, avoiding duplicate inference or tool side effects.
The proxy attributes each successful response to both the engine and requested model, keeping bounded EWMAs of end-to-end latency and completion-token throughput. Those server-owned measurements override self-reported estimates in placement snapshots, so actual service performance eventually supersedes the cold hardware prior. A multi-model vLLM engine is scored only with measurements for the model being placed; its fast model cannot lend an unrelated slow model an inflated score. For a checksum-protected managed model, every measurement is also bound to the residency's exact artifact revision. Routing ignores the previous revision immediately after a rollout, and the first successful response resets that model's estimator instead of blending incompatible revisions. Placement likewise falls back to hardware priors until revision-matching evidence exists. External vLLM inventory without artifact checksums retains the backward-compatible model-scoped behavior. Latency and throughput each ramp to full placement authority over eight relevant samples and decay against their own update timestamp. A stream that exposes no token count can refresh latency without making an old token rate look fresh. When an OpenAI-compatible stream includes final usage metadata, a bounded fragmentation-safe SSE parser extracts only its completion-token count; malformed and oversized events are ignored. Streaming responses without usage still contribute latency. The measurements are private: discovery does not expose them, managed heartbeats cannot overwrite them, and no prompt or response content is retained for allocator telemetry. Expired measurements fall back to current hardware priors until relevant new requests refresh them.
Recent ready replicas and a recently persisted demand watermark remain desired during the model's scale-down cooldown. This is the global hysteresis that prevents a quiet minute—or a signaling- server restart—from unloading a model that was just used.
Each participating computer evaluates its own telemetry independently of the controller:
- user activity and idle time;
- battery level and whether the machine is charging;
- thermal state and temperature when available;
- CPU, system memory, GPU memory, and load pressure;
- network availability;
- an explicit local drain, pause, or quarantine override.
The result is one of six lifecycle states:
| State | New work | Meaning |
|---|---|---|
accepting |
yes | Normal capacity and priority. |
throttled |
yes, reduced | The host remains useful but yields capacity. |
draining |
no | Existing work may finish before a pause or unload. |
paused |
no | The employee or battery has reclaimed the machine. |
unhealthy |
no | A confirmed safety or connectivity failure requires recovery. |
quarantined |
no | An operator has fenced the host until explicitly released. |
Debounce avoids reacting to a one-sample spike. Drain grace protects requests already running. Separate recovery thresholds and a recovery cooldown prevent rapid state flapping. Missing sensors are represented as unknown, not zero; policy may ignore unknowns or conservatively throttle.
The request router enforces the local decision too: fully accepting engines are preferred over
throttled ones, and the advertised concurrency multiplier limits new admissions. Proxy-owned active
request counters survive managed heartbeats. The server returns its per-model last_used_at
watermark to the authenticated node, which persists the monotonic value; drain and scale-down
decisions therefore reflect work actually routed by the server and survive a server restart.
Grid-owned llama.cpp children also use a durable per-host engine key. The key is stored owner-only,
sent only in an authenticated managed registration, removed from every discovery/status response,
and added by Grid on the private upstream hop. LAN clients therefore cannot bypass the routing
fence and begin new inference directly on a child port during drain. Authenticated /slots probes
account for llama work at both heartbeat and final unload boundaries.
A local override outranks global desired state. Confirmed local safety can still make an override more restrictive—for example, an operator cannot turn a critically hot machine back into an accepting one.
The planner and reconciler keep these rules even when demand, membership, or clocks change:
- Never overcommit declared memory. Reserved memory and thermal derating bound the complete desired footprint. Policy headroom additionally fences new or resized allocations, while a zero-allocation transition may retain an existing process in that margin. Unmanaged resident workloads consume capacity first.
- Never place on an ineligible host. Paused, draining, unhealthy, quarantined, stale, missing- heartbeat, or implausibly future-heartbeat hosts receive no new placement. Runtime, backend, data tier, tags, allow/deny lists, pins, and per-host model limits are enforced.
- Make capacity available before removing it. Missing desired replicas are loaded and warmed before obsolete replicas are considered for drain. The deliberate exception is an explicit higher-priority preemption on a saturated compatible host: the victim drains first, and the new model is not assigned until a later heartbeat proves the memory was released.
- Route only admitted ready models. Cached, loading, warming, and failed residencies never enter
the model identity list. A draining child may retain its identity while existing requests finish,
but the host/model admission gates remove it from active routing immediately. A managed control
envelope claiming
readyis inventory, not replacement proof: Grid reports it aswarminguntil a live, admitted child-engine record for the same host/model corroborates the route. When a corroborated child remains live but its host is intentionally draining, paused, unhealthy, or quarantined, its process state remainsreadyso reconciliation can drain and unload it; the host fence still excludes it from routing, placement, replacement, and failure- domain evidence. - Protect the last required replica. An old replica is not drained until all required desired replacements report ready.
- Drain before unload. A draining model is not unloaded while that model residency reports requests in flight; unrelated work on the same host does not block retirement forever. Managed llama ports require Grid's private engine key, so no unauthenticated direct admission can race the final idle check. If activity is unknown or exceeds the graceful deadline, non-force cleanup fails safe and leaves the proven process alive; only an explicit force stop may cut it.
- Respect ownership. Pinned, manually managed, and externally managed engines may satisfy demand but are never actuated by Grid. An unauthenticated external record also cannot authorize draining the last managed baseline replica; only authenticated managed inventory can do that.
- Bound change. Automatic mode has global and per-host concurrent-mutation limits. Minimum residency, mutation cooldown, observation timeout, and exponential failure backoff suppress churn and retry storms. Scarce execution slots go to higher-priority service even across lifecycle phases; an explicit preemption drain inherits the beneficiary's priority, while routine cleanup remains behind availability work. Within one administrator-priority class, required baseline and direct demand execute before correlation-only prewarming. Equal-priority, equal-urgency work uses its estimated remaining cold-start path, so a cached model that can serve soon is not stranded behind an unrelated artifact download. Within the same readiness class, mutation slots are filled one replica round per model, preventing one service's second replica from starting before a peer's first; capacity-release preemptions retain the same beneficiary round. Within one preemption wave, already-drained and idle capacity is released before a newly draining or busy victim. Among equally disruptive victim sets, the allocator releases a host that can start the beneficiary soonest, including cached weights and learned warm-start time; under a tight mutation budget it finishes the group with the fewest remaining lifecycle transitions instead of spending a slot on a partial release that cannot yet serve traffic. Exact-artifact loads also have an independent fanout limit (two by default). A third replica of the same SHA waits instead of joining a download stampede, while a different artifact may use the remaining global mutation budget. Pending and running loads count toward the limit, and a dependent warm is deferred with its blocked load rather than starting out of order.
- Make retries idempotent. Actions have stable IDs, pending equivalents are suppressed, and duplicate acknowledgements are harmless. Command delivery is durably marked before the response is returned to a node. A late success or failure may complete an action that the controller had cancelled; conflicting later acknowledgements are ignored.
- Fail honestly. Unmet replicas and policy shortfalls remain visible in the plan; the allocator does not invent capacity or silently relax a hard constraint.
Reconciliation separates a desired plan from side effects. Its transitions are deliberately small:
load: acquire or verify the model artifact on the selected host;warm: start it and wait for a successful readiness probe;drain: stop routing new requests to an obsolete residency;unload: release memory after the residency is drained and idle.
A warm depends on a preceding load when the artifact is not cached; if that load is deferred,
the warm is deferred too. Availability actions have priority over destructive actions. Failure
backoff is tracked per action kind, host, and model, so a broken artifact does not create a tight
fleet-wide retry loop and does not block healthy targets elsewhere. Reconciliation indexes plan
urgency, assignment memory, actual READY inventory, and mutation attempts once per tick;
safety-floor and retry construction scale with configured models, reported residencies, and retained
history rather than their cross products.
Planning likewise memoizes compatibility, capacity fit, and the exact dynamic score of a node/model/remaining-capacity/domain state within one tick. Fair replica rounds may revisit shared hosts, but repeated visits do not repeat performance, artifact-locality, hardware, or policy evaluation; colocation-enabled plans retain complete fit evaluation as their peer set changes.
If a higher service class appears while the mutation governor is full, the controller may withdraw
a lower-class constructive command only when it has never been delivered to its node. A delivered
pending command is treated as potentially running and keeps its slot until the node acknowledges
it; equal-class work is not churned merely to change queue order. When one slot is enough, a leaf
mutation is withdrawn before its useful prerequisite so reprioritization does not discard extra
work. When a host permits several
queued mutations, delivery preserves the reconciler's service ordering instead of re-sorting by
opaque action identity. A higher service class queued on a later tick also precedes an older,
undelivered lower-class entry; FIFO remains the tie-breaker within the same class.
A delivered drain or unload may already be running even while its last controller record still
says pending. If fresh placement or host evidence makes any destructive action unsafe, the
controller withdraws the entire destructive batch for that model. Every delivered member remains
listed in withdrawn_destructive, and further destructive work for that model is blocked until an
authenticated terminal receipt arrives or an authoritative heartbeat proves the action's durable
postcondition (draining/cached/absent). Availability work and destructive work for other models
continue. This guard intentionally has no wall-clock timeout: if the host never returns and never
reports a terminal receipt, destructive convergence for that model remains blocked rather than
risking a late command taking the fleet below its replacement or diversity floor.
After controller restart, restored commands receive a bounded membership-recovery grace beginning at the first reconciliation tick. If the wall clock moves backward during that grace, its in-memory anchor is rebased to the corrected time so an absent command cannot consume the mutation budget until the old future timestamp is reached.
| Mode | Plan and forecast | Proposed actions | Executable commands |
|---|---|---|---|
observe |
yes | no; drift is recorded as deferred | no |
recommend |
yes | yes | no |
automatic |
yes | yes | yes, within safety governors |
Changing away from automatic cancels pending commands. Removing a model profile creates a durable
retirement tombstone with a target of zero replicas. It cancels pending availability work, keeps
enough state to drain managed copies that reappear after an offline host returns, and clears the
model's demand history. Because local membership is not a durable inventory of every machine that
may later return, the tombstone remains visible until an operator creates that model profile again.
The initial managed backend is one Grid-owned llama.cpp child process per model. The host runtime
persists its stable host_id, local-protection state, residencies, process handle and port, latest
plan generation, and bounded action receipts. It runs only one side effect at a time while the
heartbeat remains responsive. A restart marks an interrupted action failed, proves ownership and
readiness before adopting a surviving child, and otherwise fences it. The runtime persists a child
PID and port in the immediate post-spawn callback, then enriches that record with its executable and
process-birth proof before waiting for readiness. If either durable publication fails, the child is
stopped. A confirmed-dead child releases its
handle; a live child whose exact executable, model path, alias, port, and process-birth marker cannot
be proved remains retained in failed state. That fail-closed state prevents both an unsafe signal
and a duplicate process after a transient probe failure or PID reuse.
The runtime also persists a randomly generated engine API key in the owner-only state file and
writes llama.cpp's one-key-per-line input beside it with mode 0600. It launches with
--api-key-file and --slots; the durable key never appears in process argv or the child's
environment. Readiness and activity probes authenticate with that key. Health checks snapshot
under the runtime lock, probe children in a bounded parallel pool, and commit only if the handle is
still current. Listener and port probes follow the advertised address family (0.0.0.0 for IPv4
or :: for IPv6), including IPv6-only hostnames and scoped IPv6 advertise URLs.
Commands for another host, non-executable recommendations, and older plan generations are rejected.
Dependencies must have succeeded before a command begins. A local decision that rejects admission
cancels new load or warm work, even when the global controller requested it. Unique ports are
allocated from the managed range, and the backend refuses to stop a PID it cannot prove belongs to
that model runtime. Every heartbeat refreshes physical memory and current external use. Managed
residencies are subtracted exactly once, and the node rejects a warm before process launch if its
local free-memory observation no longer satisfies the command.
An authenticated cached residency is authoritative evidence that an earlier warm lifecycle has
finished and no process remains. If later demand restores that placement, Grid may issue a fresh
warm immediately instead of waiting for the old successful WARM receipt's observation timeout;
failed warm history still retains its normal backoff. The same causal rule permits a new DRAIN when
the runtime is authoritatively ready again and a new UNLOAD when it is draining again. This lets
models cycle out and back in without a prior successful receipt imposing a false 120-second delay;
failed destructive actions remain backoff-protected.
load never infers a mutable download source from a display name. It verifies an existing GGUF, or
the authenticated profile may provide all three autonomous-transfer fields: an exact
hf://owner/repo/path.gguf source, an immutable SHA-256, and a maximum artifact size. The llama.cpp
adapter downloads under an artifact-addressed staging name, resumes bounded partial transfers,
rejects streams above the size ceiling, hashes the complete file, and atomically publishes it only
after verification. A wrong digest never replaces the prior cache. When a profile declares
artifact_sha256, both load and warm hash the exact cached file before process launch. A
residency reports the digest it proved; a same-named residency with a missing or different digest
does not satisfy placement. Grid warms a matching replica elsewhere before draining the old
version, and refuses an unsafe in-place replacement when no peer can preserve availability.
Allocator additions are namespaced so old Grid nodes can continue to register. The existing
top-level models field still means exactly "ready and routable now." A stable host_id represents
the physical machine even when several engine records run on it; the local controller merges those
records without multiplying the machine's memory capacity.
An allocator-capable node registers or updates with PUT /nodes/{node_id}:
{
"role": "allocator",
"models": [],
"host_id": "host-01HX...",
"resources": {
"capacity_mb": 65536,
"reserved_mb": 8192,
"runtimes": ["llama.cpp"],
"backends": ["metal"],
"memory_bandwidth_gbps": 400,
"compute_gflops": 27132,
"failure_domain": "floor-2",
"tags": ["employee"]
},
"allocator": {
"schema_version": 1,
"state": "accepting",
"cached_models": ["qwen3-coder"],
"residencies": [
{
"model_id": "qwen3-coder",
"memory_mb": 24576,
"state": "ready",
"loaded_at": 1785300000,
"last_used_at": 1785300100,
"managed": true,
"active_requests": 0
}
],
"actuator_capabilities": ["load", "warm", "drain", "unload"]
}
}Heartbeat requests may update load, resources, and allocator, and may acknowledge commands:
{
"node_id": "node-record-id",
"load": {"active_tasks": 2, "queue_depth": 1},
"allocator": {"schema_version": 1, "state": "accepting", "residencies": []},
"request_commands": true,
"acknowledgements": [
{
"action_id": "8ccf...",
"status": "succeeded",
"message": "ready",
"duration_seconds": 12.5,
"artifact_fetched": true
}
]
}artifact_fetched is true only when a load had to invoke the runtime's immutable-artifact
fetch path. It remains false for a cache hit, artifact verification, and every non-load action, so
the controller learns cold-fetch latency only from relevant measurements.
request_commands defaults to true for compatibility. The node sets it to false on its early
lease and fail-closed fence heartbeats: those requests update registry truth and mark placement
dirty, but return immediately without waiting for reconciliation and without durably marking any
command delivered. Only the final control heartbeat uses true and consumes returned commands.
Every allocator input advances a causal dirty revision, while only a semantic change to destructive
safety advances the separate safety revision. Repeated identical lease or relative-age telemetry can
therefore keep planning dirty without starving a command poll: availability commands may use any
successful tick at or beyond the poll's causal revision. A drain or unload is stricter. The
successful tick's safety revision must equal the current safety revision exactly, and immediately
after durably preparing the delivery marker the controller revalidates the complete destructive batch
against a fresh raw-registry snapshot under the command-selection lock. This ordering leaves no
blocking state write between the final proof and the response. Replacement control and child routes
must each have strictly more than 30 seconds left on their 60-second lease at that final check,
covering the revision wait, response delivery, and the managed node's immediate durable action-start
boundary. If a replacement becomes non-routable, falls inside that margin, or expires while the tick
or marker write is running, destructive commands remain pending and the controller durably removes
markers prepared by that poll before suppressing the response. A failed compensating write retains
the marker conservatively and exposes the uncertainty in allocator status; availability work can
still proceed.
An authenticated heartbeat response carries commands for that host_id:
{
"ttl_seconds": 60,
"model_last_used_at": {"qwen3-coder": 1785300100},
"allocator": {
"mode": "automatic",
"controller_lease_ttl_seconds": 14.8,
"commands": [
{
"action_id": "8ccf...",
"kind": "warm",
"node_id": "host-01HX...",
"model_id": "qwen3-coder",
"memory_mb": 24576,
"plan_generation": "c91d...:00000000000000000042:4fc1...",
"controller_term": 7,
"controller_id": "c57a...",
"controller_lease_expires_at": 1788020000.0,
"dependencies": [],
"executable": true
}
]
}
}The plan generation is a persistent epoch plus a monotonically increasing sequence and an input
digest, so plans remain ordered across wall-clock changes. Mutation authority is a separate durable
fence: automatic mode acquires a renewable single-writer lease beside the controller state file,
increments its term on every takeover, and stamps every command with (controller_term, controller_id, controller_lease_expires_at). At each authenticated command response, the server
also converts that absolute authority deadline into a remaining TTL in its own clock domain. The
node subtracts the complete request duration and admits the command using only its local monotonic
clock; a fast or slow node wall clock therefore cannot extend authority or reject a valid leader.
A managed node durably remembers the highest term, accepts at most one controller identity in that
term, rejects expired TTLs, and rejects every lower-term command even after restart. The absolute
field remains for controller persistence and backward-compatible peers. This makes command safety
independent of network delivery order and cross-host wall-clock skew; plan generations continue to
order plans within the accepted authority.
The complete action also includes its reason, creation time, and not_before time. Wire objects use
schema_version: 1; the new authority fields are additive so older persisted actions still decode,
but once a node has observed a fenced command it will not return to the legacy term-zero namespace.
Unknown or malformed residency rows are
excluded and make the host snapshot unhealthy; an invalid host lifecycle likewise fails closed
rather than becoming eligible. Managed registry IDs are derived from the authenticated host and
model identities, so a host-scoped credential cannot squat another host's control or engine record.
Managed engine registration additionally carries a private engine_api_key; the signaling server
stores it only in memory for upstream forwarding and never serializes it into public node output.
An HTTPS engine may also carry a bounded CA chain inside the authenticated allocator envelope.
Grid removes that PEM before storing public metadata and builds a private hostname-verifying SSL
context. Residencies carry model-local loaded/last-used ages; Grid reconstructs timestamps from its
receipt time so node clock skew cannot bypass minimum-residency or scale-down cooldown.
A remote allocator deployment has one controller and any number of managed provider nodes. Run the controller on the same machine as the relay and bind it only to loopback; the relay is the sole authenticated bridge between remote nodes and the controller. Requests continue to enter the normal relay/router path. The relay sends bounded lifecycle features to the allocator, and the allocator sends placement commands to enrolled nodes. Neither prompts nor responses are forwarded to or retained by the allocator.
On the controller/relay machine, start a dedicated local Grid runtime for allocator state and API:
grid --local start allocator-control \
--host 127.0.0.1 \
--advertise-host 127.0.0.1 \
--port 22101Find its generated config under ~/.grid/grids/<allocator-grid-id>/config.json. Keep that file
owner-only. On a self-hosted relay using autonomous-grid-cli, configure the relay master with the
supported environment command; use the literal absolute config path rather than ~:
grid network set-env <remote-grid-id> \
GRID_ALLOCATOR_SIDECAR_URL=http://127.0.0.1:22101 \
GRID_ALLOCATOR_ENROLLMENT_TOKEN_FILE=/absolute/path/.grid/grids/<allocator-grid-id>/config.json
grid network restart-server <remote-grid-id>For a relay managed by another service runner, set those same two environment variables on that
relay process. GRID_ALLOCATOR_SIDECAR_URL accepts only literal loopback HTTP. The enrollment file
is read locally by the relay and its operator capability is never returned to a provider.
Each capacity machine needs a live remote provider identity, then enrolls that same authenticated identity as allocator-managed capacity. An already-serving node only needs the allocator command. A fresh capacity node can create an empty provider explicitly; it does not need a fake bootstrap model, and it advertises no inference route until the allocator has loaded one:
grid --remote sync
grid --remote use <grid>
# Fresh node with no engine yet:
grid --remote join <grid> --allocator-provider --name <node-name>
grid --remote allocator join <grid> --dedicatedEnrollment verifies the llama.cpp runtime before advertising the node's managed capabilities. On a fresh machine it installs Grid's version- and SHA-256-pinned build automatically. It also retries the narrow provider-registration race for up to 15 seconds, so these commands may be run back to back; authentication, policy, and unrelated conflict failures are never retried.
After installing a newer Grid build on a provider, apply it without manually locating the detached daemon or its controller-sidecar scope:
grid --remote allocator join <grid> --dedicated --restartRestart uses the normal route fence and request drain, stops only identity-proven allocator-owned children, obtains a fresh host-scoped credential, then adopts cached state under the same provider identity. Other provider routes and the relay stay online.
Use --dedicated only for an always-on server. Omit it on a workstation or laptop so local activity,
battery, thermal, memory, disk, and network protection can throttle or fence allocator work. The
enrollment response contains only a host-scoped credential, retained by the detached node process;
operators never copy the controller capability to workers.
Policy administration remains local to the controller in this release. Register model profiles and move through the rollout modes from the controller machine:
grid --local allocator model set <model.gguf> \
--grid allocator-control \
--memory-mb <resident-mb> \
--artifact-sha256 <64-hex-digest> \
--artifact-source hf://owner/repo/path/to/model.gguf \
--artifact-size-mb <download-mb> \
--runtime llama.cpp \
--min-replicas 0 \
--max-replicas 3
grid --local allocator mode observe --grid allocator-control
grid --local allocator tick --grid allocator-control
grid --local allocator status --grid allocator-control
grid --local allocator mode recommend --grid allocator-control
grid --local allocator status --grid allocator-control --json
grid --local allocator mode automatic --grid allocator-controlVerify both control-plane convergence and real serving:
grid --local allocator status --grid allocator-control
grid --remote models <grid>
grid --remote chat -m <model.gguf> "hello"Before enabling a newly installed engine adapter on a physical provider, run its real lifecycle qualification locally on that host. It validates inventory, immutable identity, warm, ownership, native readiness, real inference, activity, and drain/stop, then writes durable evidence:
grid --local allocator qualify ollama <model> --artifact-sha256 <digest>
grid --local allocator qualify comfyui comfyui:image_generation
grid --local allocator qualify vllm <model> \
--artifact-source hf://owner/repo@<commit> \
--artifact-sha256 <snapshot-identity> --artifact-size-mb <bound>See physical runtime qualification for exact behavior.
On a remote provider enrolled with --dedicated, the node uses Grid's multi-engine lifecycle
orchestrator. It manages installed llama.cpp, Ollama, ComfyUI, and vLLM runtimes through one
LOAD → WARM → READY → DRAIN → UNLOAD contract while the existing provider remains the only relay
ingress. Ollama and ComfyUI stay bound to loopback and are advertised only after their native
readiness APIs prove the requested model or workflow is usable. Ollama downloads require an exact
ollama://<model> source plus the registry digest; vLLM downloads require an immutable
hf://<owner>/<repo>@<revision> snapshot identity. ComfyUI workflow assets must already be
installed because a workflow manifest—not a mutable model name—is the safe unit of deployment.
Enrollment grants lifecycle authority only to processes started or process-identity-proven by the allocator. A pre-existing external route remains visible demand and capacity evidence; Grid never kills an unrelated daemon merely because its API resembles vLLM or Ollama. To migrate such a route, first configure a versioned allocator model profile and prove its canary, then retire the old external route after the managed replacement is READY.
Start the local grid and join each computer that should offer managed capacity. Pre-pulling remains supported and avoids cold network transfer:
grid up
grid pull <hugging-face-repo>:<model.gguf>
grid allocator node start
grid allocator node statusFor an already-serving provider on a remote Grid, enrollment is one command and reuses its existing Grid membership:
grid allocator join forgeThe relay verifies that the token-bound node is currently a provider, derives its stable allocator host identity, and obtains a host-scoped credential from the controller. The operator capability is never sent to the worker, and the returned node credential remains in memory rather than appearing in the terminal or process arguments.
Create a placement profile from a machine that can control the grid. Memory is the resident runtime budget for one replica, not the file's compressed size:
grid allocator model set <model.gguf> \
--memory-mb 12000 \
--artifact-sha256 <64-hex-digest> \
--artifact-source hf://owner/repo/path/to/model.gguf \
--artifact-size-mb 9000 \
--workload-score coding=1 \
--workload-score research=.8 \
--max-colocated-models 1 \
--colocation-exclude <interfering-model.gguf> \
--min-replicas 1 \
--max-replicas 3 \
--min-failure-domains 2The profile command also accepts repeated --runtime, --runtime-memory-mb RUNTIME=MB, --backend, --required-tag,
--forbidden-tag, --pin, and --workload-score WORKLOAD=SCORE values. Workload scores are
capability hints in (0, 1] for portfolio planning; they do not route an individual request. Data
tier, target utilization, expected service time, latency SLO, priority, load/warm estimates,
residency and scale-down cooldowns are explicit flags.
--artifact-sha256 is optional but recommended for managed production GGUFs; it is canonicalized
to lowercase and becomes part of command, retry, and readiness identity. --artifact-source
enables autonomous fetch and therefore requires both the digest and --artifact-size-mb; sources
without all three fields are rejected. The source URI is operator configuration, not a place for
embedded credentials.
--max-colocated-models is also optional (0 means unlimited). The value counts the candidate
itself, so 1 requests exclusive serving for an interference-sensitive model. The constraint is
reciprocal: Grid will neither place that model beside another live/planned model nor later place a
different configured model beside it. Cached-only weights do not consume a serving slot. When a
managed host already violates a tightened ceiling, Grid deterministically elects the higher-priority
(then more constrained) survivor and stages safe drain/unload of removable peers before it admits
new work. Existing manually managed inventory that violates a profile remains visible but is
reported unsatisfied; the constraint never grants Grid authority to resize or stop an external
engine. When several compatible runtimes are installed, a new placement chooses the runtime with
the smallest declared resident footprint (then a stable name tie-break). A live compatible
residency is sticky: changing engines is a versioned canary/replacement operation, never an in-place
restart.
For a narrower policy, repeat --colocation-exclude <model> to name only measured bad pairings.
Exclusions are reciprocal even if declared by one profile: neither placement order can put the pair
together. Compatible peers may still share the host, and the same managed-only staged convergence
applies if a pair is already live when the policy is added.
--replica-concurrency declares a conservative service-slot estimate for a newly managed replica.
Once a single-model engine is ready, its live max_concurrency may prove a higher batch width; a
multi-model engine's shared node-wide limit is never credited independently to every model. Queue,
latency, or error pressure still requests at least one replica beyond the current ready set.
If --runtime is omitted, it defaults to llama.cpp. Once the flag is present, only the listed
runtimes are eligible.
Use the three modes as a rollout sequence:
grid allocator mode observe
grid allocator tick
grid allocator status
grid allocator mode recommend
grid allocator tick
grid allocator status --json
grid allocator mode automaticrecommend is the default. Before selecting automatic, either pre-pull the exact profiled GGUF on
eligible managed nodes or configure its exact source, digest, and size ceiling. An uncached model
without an approved source fails safely and enters backoff rather than downloading inferred
weights.
status shows host lifecycle and capacity, model count, pending mutations, withdrawn destructive
commands, unmet constraints, the dirty/processed/success/safety revisions, and any persistent-state
warning; --json includes the full snapshots, demand forecasts, plan, reconciliation result,
pending_commands, delivered_pending_action_ids, withdrawn_destructive, the latest bounded
delivery-safety error, and bounded history. tick is useful after a profile or host change. The
server also runs a periodic pass and coalesces registration, heartbeat, acknowledgement, and demand
events into prompt background passes without blocking request serving.
To retire a profile or this machine's managed node:
grid allocator model remove <model.gguf> # `rm` is an alias
grid allocator node stopnode stop first requests a graceful local drain, waits for active requests, and then stops owned
model processes. Startup failure and a stuck runtime use bounded escalation against the verified
detached process group or Windows process tree. The
daemon advertises only children that pass a steady health check, retries failed registry deletion,
and uses an instance-scoped readiness lease so stale state cannot make node start report success.
Before any signal, the CLI verifies the daemon's unique command-line instance and process-birth
marker; an ambiguous, legacy, or reused PID is never killed automatically and requires manual
inspection. If a node credential expires, its children keep
serving until the server-side routing lease has expired, avoiding a stale route to a dead port.
Local operators can fence a machine independently of the global controller; these overrides are
durable across node restarts and may expire automatically:
grid allocator node drain --reason "taking laptop home" --for-seconds 3600
grid allocator node pause --reason "battery use"
grid allocator node quarantine --reason "investigating thermal fault"
grid allocator node resumeEvery command accepts --grid <name|id|local-url> at its own subcommand level. Administrative
commands use the operator capability from the local grid config, GRID_ALLOCATOR_CONTROL_TOKEN, or
their --token-file. A managed node uses a separate host-scoped credential from
GRID_ALLOCATOR_NODE_TOKEN or node start --token-file. Never copy the operator capability to a
worker. The allocator CLI is local-mode only in this release.
To provision another computer without putting the capability in terminal output or shell history, write it to an owner-only file on the controller and transfer that file over your existing secure administration channel:
grid allocator token write ./grid-node-token --host-id host-mac-studio
# securely copy the file to the other computer, then:
grid allocator node start --grid https://grid.company.internal \
--token-file ./grid-node-token \
--advertise-host worker-01.company.internal \
--engine-tls-cert ./worker-01-chain.pem \
--engine-tls-key ./worker-01-key.pem \
--engine-tls-ca ./company-inference-ca.pemtoken write signs an expiring credential authorized only for the selected stable host ID; if the
ID is omitted, it prints the generated ID once for provisioning. The file is created with mode
0600 on POSIX and an owner-only ACL on Windows. The secret is never stored in the node process
record, command line, public node metadata, model-child environment, or CLI output. The TLS private
key must be owner-only. The certificate SAN must cover the exact advertised hostname or IP, and
--engine-tls-ca supplies a private intranet CA to the node and Grid upstream verifier. Grid
refuses to send node or engine credentials over non-loopback plain HTTP; --allow-insecure-http
is accepted for CLI compatibility but does not override that boundary for managed nodes. A node
started against a Grid owned by the same machine advertises the Grid's literal loopback control
address by default. Remote workers must use HTTPS for Grid control and TLS for their advertised
engine address. grid down stops a managed
allocator node before it stops the local signaling server, allowing owned model processes to
unregister and exit cleanly.
The local signaling server exposes:
GET /allocator/status— current mode, host snapshots, profiles, forecasts, latest plan, reconciliation result, last successful tick duration, pending commands, withdrawn destructive commands, and bounded action history;PUT /allocator/models/{model_id}— create or replace a model profile;DELETE /allocator/models/{model_id}— retire a model profile and safely converge to zero;PUT /allocator/mode— selectobserve,recommend, orautomatic;POST /allocator/tick— request an immediate reconciliation pass.
Administrative routes require the durable operator capability in X-Grid-Allocator-Token or a
Bearer header. Node registration, heartbeat, command delivery, and acknowledgements require the
host-scoped credential in X-Grid-Allocator-Node-Token; the server rejects use against another
host ID. Neither credential belongs in engine metadata or logs. Read-only status follows the local
server's existing LAN visibility.
Operational rollout should follow this order:
- Register hosts and verify stable physical
host_idvalues, capacity, compatibility, policy tags, actual residencies, and heartbeat freshness. - Add model profiles with explicit memory and replica bounds.
- Run in
observe, thenrecommend, and inspect unsatisfied constraints and proposed mutations. - Verify local protection transitions on employee machines and confirm external engines appear as manually managed.
- Select
automaticonly after the proposed placements and mutation limits match the fleet's failure tolerance. Returning torecommendis the kill switch for new automatic work.
Controller state is written atomically. If that file is corrupt at startup, Grid quarantines it,
starts a clean controller in recommend mode, and keeps a visible warning in allocator status.
If no durable state path exists, or the requested path cannot be quarantined or written,
automatic mode is refused rather than running mutations with non-durable intent.
After a valid state restore, fresh membership must re-register before destructive work resumes;
the restart grace period prevents a temporarily incomplete fleet view from causing unloads.
Use the deterministic scenario lab to explore a large heterogeneous fleet without starting model processes or pretending the development Mac owns the modeled GPUs:
uv run grid test scenario \
--machines 8 \
--models 8 \
--users 50 \
--duration 30m \
--seed 42
uv run grid test scenario --machines 4 --models 8 --users 50 --duration 30m \
--seed 7 --oracle
uv run grid test scenario --machines 16 --models 9 --users 500 --duration 2h --json
uv run grid test scenario --machines 8 --models 8 --users 50 --duration 2h \
--workload-trace coding=/path/to/trace.csv \
--workload-trace image=/path/to/image-trace.csv
uv run grid test scenario --machines 4 --models 8 --users 50 --duration 2h \
--seed 144 --strategy greedy
uv run grid test graduate --machines 2,4,8 --seeds 42,144 --duration 2hThe lab creates logical Apple/Metal, ComfyUI/MPS, and NVIDIA/vLLM/ComfyUI configurations with different memory, disk, cached artifacts, concurrency, and performance. User personas produce coding, research, marketing, sales, design, image, video, embedding, and general demand through the real bounded request classifier; operations traffic also names the baseline model so direct demand and autonomous portfolio demand compete in the same run. A seeded workday includes a coding surge, creative campaign, thermal throttle, node outage, recovery, and cooldown. Every planning tick uses the production workload intelligence and placement planner. Personas are grouped into bounded synthetic projects whose request timestamps follow stable work stages while request counts remain independently sampled. The production learner must discover those sequences from observations; the scenario never inserts a forecast directly.
The report explains joint portfolio changes, loads, unloads, node transitions, capacity shortfalls,
persistent workload-admission states, demand served, least-user service and user/workload SLO
attainment, portfolio suitability, memory use, cache locality,
persistent modeled disk consumption, cold starts, capacity recommendations, and
safety invariants. Requests are scored against the model that was actually READY and selected when
they arrived. When routed device-time exceeds that model's effective capacity, the lab feeds the
resulting deterministic failures, latency inflation, and queue depth into the next production
controller tick; the timeline and summary report the overloaded models instead of letting a
resident-but-saturated engine masquerade as successful service. Admission metrics separately count
state and concrete blocking-model minutes, so
an unselected workload cannot masquerade as a low replica shortfall. It intentionally reports
shortfalls instead of inventing capacity. --timeline
prints every changing tick; --json emits the complete stable report;
reusing --seed reproduces the same run. Artifact disk constraints are translated into each
one-model logical node's admission set, while the allocator's native runtime, backend, lifecycle,
memory, headroom, and model-slot rules remain authoritative.
The simulator also applies the reconciler's minimum-residency fence before materializing a desired
removal. Deferred drains remain physically resident and appear in the timeline, so an aggressive
plan cannot earn fake service or churn credit for a transition the real node would refuse.
The overall score keeps raw served demand, suitability, SLO attainment, lifecycle efficiency, and
safety, but evaluates minimum-user and SLO terms over structurally allocatable workloads. A
runtime/backend-incompatible video request still lowers raw fleet service and appears as a capacity
gap; it does not make every allocator tie at zero on decision quality. A normalized logarithmic
feasible-workload coverage term makes abandoning an entire supported user need visible while still
rewarding improvements near full service.
--strategy runs the exact same topology, user population, arrival trace, failures, routing
physics, and one-tick startup delay under one of four controls. smart is the production workload
intelligence and planner. reactive removes workflow correlation, lookahead, and predictive cache
work. greedy maps only the current minute's visible workload to its highest-suitability feasible
model and uses the ordinary constrained placement planner. static receives the declared user
population's average workload mix once, fixes that host/model map, and never reschedules around a
failure. These are deliberately competent baselines; they do not receive future request arrivals.
grid test graduate runs all four strategies for every requested fleet size and seed. It refuses
to graduate on a single attractive score: the machine-readable gates require zero smart-policy
safety violations, identical demand traces, bounded score and service regressions, a material
median service gain over static placement, substantially less churn than minute-greedy control,
competitive least-workload service on the largest fleet, and measured predictive startup savings
over reactive-only control. A failed gate is the work queue for the next allocator improvement;
passing the matrix is the feature-freeze criterion for the logical-machine phase.
Scenario artifacts carry deterministic immutable identities and sources. A learned correlation-only
prefetch therefore consumes modeled disk without consuming a model slot, and a later placement
records a real cache hit. The scorecard reports prefetch downloads, hits, unused predictions,
expiration/reclaimed disk, lead time, and avoided load seconds so a predictor that merely fills
disk cannot look successful.
Managed nodes retain provenance for artifacts they actually downloaded because of a predictive
prefetch. If such an artifact is never warmed and demand disappears, the planner may expire it after
the predictive cache TTL (six hours by default). Expiration is exact-SHA fenced and applies only to
an unused managed predictive cache entry; existing operator caches, pinned artifacts, active models,
used downloads, manually managed nodes, and older nodes without the evict actuator remain
untouched. Queued eviction commands are revalidated against the latest heartbeat before delivery.
When several workflow predictions compete for the bounded prefetch budget, Grid ranks them by
confidence- and pressure-weighted cold-start value per artifact MB. The estimate uses current
predicted concurrency and learned per-host transfer time, then is recomputed after every chosen
transfer because each choice changes the remaining node disk ledger.
Repeated within-session workflow transitions also learn a bounded median stage-to-stage lead time.
When that timing is known, cache value counts only transfer work that can finish before predicted
demand; an enormous download with five seconds of lead no longer beats a smaller artifact that can
remove more of the actual cold-path wait. Unknown timing retains the ordinary load-time estimate.
If authenticated disk is full, a mature unused predictive artifact may be replaced by a stronger
prediction. The incoming artifact must provide at least the configured value-per-byte gain (2× by
default), must not have less total expected startup value, and the victim must have spent at least
15 minutes in cache. Direct or queued demand, live/used weights, pins, operator-managed artifacts,
and inexact revisions are never replacement candidates. Replacement is a two-generation
transition: Grid evicts the exact victim first, waits for a heartbeat proving disk was reclaimed,
then plans the incoming prefetch. The planner never credits intended deletion as free space.
When fragmented free space requires several smaller victims, Grid searches a bounded union of the
lowest-value and largest mature predictive artifacts (up to three victims by default). It proves
the incoming model beats the complete group's lost value and aggregate value density before
issuing even the first eviction. The ordinary mutation budget may execute that proven group over
several heartbeats; if demand changes, the remaining directives are replanned rather than blindly
continued.
--oracle adds a bounded exhaustive benchmark for at most four machines, nine models, and 240
minutes. It replays the exact observed request trace with perfect future knowledge, exhaustively
scores every runtime/backend/memory/node-state-compatible placement, enforces a two-minute-or-longer
hold and one-minute load-to-ready timing, and minimizes mutations among service-maximizing
schedules within a 256-state-per-block tie-search cap. The report shows the service ceiling,
remaining potential gain, winning placement
changes, states evaluated, and a cumulative artifact-disk audit. This is a development target, not
a production policy: perfect foresight and optimal routing are intentionally optimistic, and the
scale-down cooldown is treated as a tunable policy rather than a hard schedule fence. If artifact
and startup audits pass, the measured gap is a hindsight placement opportunity independent of disk
and startup limits; otherwise the report states that cache preparation is also required.
For production-shaped replay, --workload-trace WORKLOAD=CSV accepts a headerless time series with
timestamp seconds and request rate in its first two columns; additional distribution columns are
ignored. Repeat the option for independent coding, research, image, video, or other supported
workloads. Grid maps the whole trace onto the requested scenario duration and normalizes its mean
to the built-in workload curve. This changes burst timing while keeping capacity comparisons on the
same offered-demand scale. Trace input is bounded to 5 MiB and 100,000 rows per workload and is
validated for finite, nonnegative rates and strictly increasing timestamps.
The scenario lab is a planning-scale and decision-quality test only. It is not an inference test. The persistent fixture below is the real-process proof: every successful text or image result comes from an engine running on the development Mac.
For interactive development, start a persistent Grid with any number of logical machines. Each machine gets a stable host id, failure domain, state file, credential, capacity share, and real llama.cpp child while the Grid API remains available until explicitly stopped:
uv run grid test start --machines 4
uv run grid test status
uv run grid test watch
uv run grid test demo --users 6 --requests 12
uv run grid --local models http://127.0.0.1:22100
uv run grid --local chat --grid http://127.0.0.1:22100 \
-m SmolLM2-135M-Instruct-Q3_K_M.gguf 'Reply with OK'
curl http://127.0.0.1:22100/v1/models
curl http://127.0.0.1:22100/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{"model":"SmolLM2-135M-Instruct-Q3_K_M.gguf","messages":[{"role":"user","content":"Reply with OK"}],"max_tokens":8}'
uv run grid test stopFor a real mixed-framework Grid, install the runtime and bundle once, then reserve one of the N logical machines for ComfyUI. On Apple Silicon this is ComfyUI with PyTorch/MPS; the remaining logical machines run independently managed llama.cpp/Metal children:
uv run grid engine install comfyui
uv run grid engine pull z_image
uv run grid test start --machines 4 --include-comfyui --media-bundle z_image
uv run grid test demo --users 6 --requests 12 --max-tokens 32Here --machines 4 means four total logical machines: three llama.cpp text nodes and one ComfyUI
media node. All processes still share one physical Mac, so the fixture partitions reported capacity
instead of multiplying it. It does not pretend that CUDA or vLLM exists on Apple hardware. A CUDA
host can install/register vLLM as external inventory during the physical-node phase.
--machines N accepts 1–32 logical machines; practical limits are the physical machine's memory
and process capacity. Use --model, --port, and --engine-port-base to run a different cached
GGUF or avoid local port conflicts. Starting is idempotent for matching settings, and status can be
emitted as JSON. watch follows residency transitions and allocator command outcomes without
stopping the Grid when you press Ctrl-C. The ordinary shipped CLI can address the test endpoint by
URL (--local makes that explicit even when remote mode is active), including grid chat,
grid models, and grid allocator status --grid http://127.0.0.1:22100. The start output names the
owner-only token file for allocator mutations. The fixture stays isolated from the active local or
remote Grid configuration.
To compare genuinely different models and heterogeneous node capacities, add repeatable portfolio candidates plus one capacity per text node, then run the real competition:
uv run grid test start --machines 4 --include-comfyui --media-bundle z_image \
--portfolio-model qwen2.5-coder-0.5b-instruct-q4_k_m.gguf \
--candidate-model qwen2.5-0.5b-instruct-q4_k_m.gguf \
--workload-model coding=qwen2.5-coder-0.5b-instruct-q4_k_m.gguf \
--workload-model research=qwen2.5-0.5b-instruct-q4_k_m.gguf \
--workload-model marketing=qwen2.5-0.5b-instruct-q4_k_m.gguf \
--workload-model sales=qwen2.5-0.5b-instruct-q4_k_m.gguf \
--text-capacities-gib 3.5,5,25.5
uv run grid test demo --users 6 --requests 12 --max-tokens 24
uv run grid test competeWith explicit workload bindings, grid test demo is a real adaptive workday rather than a static
smoke test. It starts from one tiny baseline, observes four user workload classes, and requires the
allocator to load a coding specialist and one shared research/marketing/sales model on the two
larger hosts. A real general-demand surge must then offload both specialists and scale the baseline
from one to three replicas. When that surge expires and only research/marketing remain, Grid must
return to one baseline plus the shared non-coding model while leaving coding capacity off. Final
cooldown must drain/unload every optional model. The command prints every residency transition and
requires genuine OpenAI-compatible responses at each mix; --include-comfyui additionally requires
a real generated PNG. Forecast history is retained between runs, so the demo deliberately converges
to idle and freezes actuation before admitting the next request mix instead of clearing demand.
grid test compete loads one candidate at a time through the production lifecycle, runs eight
deterministic coding questions through real llama.cpp/Metal inference, and submits authenticated
correctness and latency evidence without manufacturing user demand. It then offloads every
candidate and sends unresolved coding traffic through the ordinary Grid request boundary. The
allocator—not the router—chooses the measured portfolio winner, reloads it, verifies a real answer,
and fails the test if it did not use the planner-preferred capable logical node. It then makes the
winner fleet-ineligible with a temporary impossible node tag, verifies that the allocator explains
the rejection, unloads it, loads a feasible runner-up, and serves another real answer. Removing the
constraint must restore the measured winner. Evaluation-marked
inference still updates per-engine performance, but only the owner capability may mark it and its
quality is recorded separately at POST /allocator/evaluations; ordinary callers cannot suppress
their demand signal. Evaluation submissions may include artifact_sha256; an explicit digest must
match the currently configured revision. The command leaves the selected winner ready for
interactive grid chat use.
Configured logical capacities must fit within this machine's real usable memory. The list length
must equal the number of text nodes (--machines minus the optional ComfyUI node), so incompatible,
constrained, and flexible hosts can be tested explicitly.
Without --workload-model bindings, grid test demo performs the original focused coding flow. It
uses no synthetic inference or demand injection. It first converges the
baseline to one real replica and verifies a genuine OpenAI-compatible response. Three distinct
users send genuine coding requests for unresolved auto and receive honest HTTP 503 responses
with no router involved.
The production request path classifies bounded features locally; after its minimum evidence
threshold, the allocator projects that workload onto the configured coding portfolio and
proactively loads and warms a real llama.cpp canary before any request names it. The first named
specialist call must then return real output. Multiple client personas send concurrent requests to
both text models with stable affinity keys and bounded production-style retries. The report requires
non-empty assistant output, response IDs, usage, client-visible latency including retries, per-node performance samples,
and the complete process lifecycle before passing. Observed demand expires naturally; the command
never clears or fabricates it.
When the Grid was started with --include-comfyui, the same command also sends a genuine image
workflow through /v1/media/image/generate, validates returned PNG bytes, and writes the result into
the logical test run directory. ComfyUI is currently registered as immutable media inventory: the
fixture owns its process startup and teardown, while allocator mutations remain limited to
Grid-owned llama.cpp residencies. This ownership boundary is reported rather than hidden.
Before a physical multi-machine rollout, the development harness can partition one Mac into isolated logical hosts with separate host IDs, failure domains, durable state files, credentials, port ranges, capacity shares, and real llama.cpp children:
uv run python tests/e2e_allocator_logical.py --nodes 2
uv run python tests/e2e_allocator_logical.py --nodes 4
uv run python tests/e2e_allocator_logical.py --nodes 4 --scenario activity
uv run python tests/e2e_allocator_logical.py --nodes 2 --scenario restart
uv run python tests/e2e_allocator_logical.py --nodes 4 --scenario preemption \
--second-model <cached-alias.gguf>The lifecycle runs cover demand-driven warm, real OpenAI-compatible inference, route fencing,
drain, unload, and abrupt child recovery. The activity scenario runs three replicas with a fourth
logical spare, marks a loaded partition user-active, and requires make-before-break evacuation.
The restart scenario rebuilds each logical node agent from durable state and requires it to adopt
the exact live llama.cpp PID before proving that the adopted process can still drain and unload.
--scenario contention --second-model <cached-alias.gguf> exercises two model identities under a
one-model-per-logical-host claimed capacity budget. Logical performance and memory telemetry are
partitioned so the harness never reports N times the physical Mac's capacity.
--scenario preemption --second-model <cached-alias.gguf> keeps demand for a low-priority model
active on every logical host, injects a high-priority burst for the second model, and requires real
drain/unload of every incumbent before the second model is warmed and served. It then retires the
critical burst and requires the displaced batch service to warm back onto every logical host and
complete another real streamed request before final cleanup.
The design follows several primary systems results while preserving Grid's allocator/router split:
- Scalable Joint Resource Allocation for SLO-Constrained LLM Inference in Heterogeneous GPU Clouds motivates joint feasibility, model choice, provisioning, routing, quality, latency, and memory constraints. Grid applies its fleet-feasibility and model portfolio insight in the allocator while leaving per-request routing independent.
- Llumnix and Libra motivate dynamic rescheduling, isolation, and SLO-aware adaptation under changing load. Grid's load/warm and drain/unload state machine applies those ideas at the slower fleet-allocation timescale.
- Strata identifies cache- loading latency, fragmented storage, and concurrent delay hits as first-class scheduling costs. Grid applies the portable parts at whole-model scope: learned load time, deadline-aware prefetching, aggregate fragmented-cache replacement proofs, and bounded exact-artifact fanout.
- Remote allocator nodes require a relay with the authenticated allocator sidecar and enrollment bridge enabled. Policy administration remains controller-only; remote providers may enroll only their own already-live identity.
- Dedicated allocator providers have lifecycle adapters for Grid-owned llama.cpp and vLLM children, plus model-memory control for loopback Ollama and ComfyUI services. External/manual processes remain inventory and routing sources until explicitly enrolled; lifecycle authority never follows from protocol detection alone. LM Studio and generic API engines remain routing-only.
- The current autonomous.ai NVIDIA engines are vLLM/CUDA even though live discovery labels their
ownership class
external. Framework identity and lifecycle ownership are independent: those engines participate in routing and placement evidence, but discovery alone does not grant Grid permission to start, drain, or stop them. Local auto-discovery publishes the detected runtime; when pointing at an engine explicitly, usegrid join --at <url> -m <model> --kind vllm(or the corresponding kind) so runtime-constrained profiles can use the inventory. - Managed llama.cpp and vLLM can autonomously fetch exact, size-bounded immutable artifacts. Ollama verifies the native registry digest and refuses to overwrite a pre-existing tag with another digest. ComfyUI workflow assets still require their runtime-specific installation path and are unloaded from accelerator memory rather than deleted.
- Capacity is refreshed by the node as stable physical capacity plus dynamic non-Grid reserve.
Device count and per-device VRAM are preserved, and profiles may fail closed with
min_gpu_countandmin_gpu_memory_mbconstraints. This covers basic tensor-parallel feasibility, artifact disk admission, and safe predictive-cache eviction but does not yet model GPU interconnect bandwidth, heterogeneous sharding, NUMA boundaries, or simultaneous transfer- bandwidth bottlenecks. - Model profiles accept a portable
memory_mbfallback plus runtime-specificruntime_memory_mbestimates, so llama.cpp/Metal and vLLM/CUDA placements account for their distinct footprints. If a node advertises several matching runtimes, the planner conservatively uses the largest matching estimate. Managed GGUF profiles can additionally require an immutable SHA-256 and operator-authorized exact artifact source; source-revision resolution remains operator-managed rather than silently following mutable repository state. - The planner is a transparent deterministic heuristic, not an optimal mixed-integer solver. It prioritizes predictable safety and understandable decisions over a mathematically minimal cost.
- Inter-model interference is controlled through the per-profile
max_colocated_modelsceiling and explicit reciprocalcolocation_excludespairs. Grid does not yet infer those pairs from production co-run experiments or partition GPU execution resources such as CUDA MPS/MIG. - In-memory LAN node membership is rebuilt by registration after a local signaling-server restart; durable controller state does not make a stale node eligible without a fresh heartbeat.