You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
fix(agents): make deployed agent inference URL container-reachable
Deployed agents baked the API pod's own base URL into their NAT config
llms.*.base_url via inject_gateway_url() at API-create time. Under
embedded-PDP auth the API pod's NMP_BASE_URL is a loopback address
(http://localhost:8080), and in k8s an agent pod resolves loopback to
itself, so agent model calls never reach the platform. If the platform
base URL instead pointed at an auth front-door proxy, agents looped
forever on 503s and could DoS it (and everything behind it).
The only mode-aware rewrite (container_gateway_url) fed a dead
NMP_GATEWAY_BASE_URL env var that nothing reads, so neither the docker
loopback rewrite nor any k8s fix ever reached the config the agent runs.
This fix hands an agent only a base URL we know is container-reachable,
never the raw platform base URL.
- resolve_agent_gateway_url() returns a known-good target per mode:
k8s uses the in-cluster API Service DNS; docker rewrites loopback
(including IPv6 [::1]) to host.docker.internal and passes other hosts
through. It raises for unsupported modes or when k8s has no internal
URL, rather than deploying an agent that cannot reach the platform.
- rewrite_config_base_urls() rebases each Inference Gateway
llms.*.base_url onto that reachable address, preserving the path.
- get_internal_base_url() reads NEMO_INTERNAL_BASE_URL then
NMP_INTERNAL_BASE_URL; also agents.deployments.k8s_internal_base_url.
- Removes the dead NMP_GATEWAY_BASE_URL env var.
- Helm sets NMP_INTERNAL_BASE_URL to the internal API Service DNS on the
api and controller pods.
- Updates deploy-agents docs; adds unit tests for docker/k8s resolution,
the fail-fast path, and config rebasing.
Signed-off-by: Ben McCown <bmccown@nvidia.com>
0 commit comments