System Info
verl-omni GitHub Actions GPU smoke workflow on the dynamically provisioned L20x8 runner.
Observed on PR #524:
Information
Tasks
Reproduction
- Add the
ready-for-ci label to a pull request so that the GPU smoke workflow starts.
- While the GPU smoke job is still running, update or force-push the PR branch.
- Trigger
ready-for-ci again for the updated revision.
- The previous workflow is cancelled through the workflow concurrency group, and a new dynamic GPU runner job is scheduled.
- During the new GPU smoke run, a diffusion worker may fail to initialize its
torch.distributed.TCPStore because the selected port is already in use.
The first hard error from the affected run was:
torch.distributed.DistNetworkError: The server socket has failed to listen on any local network address. port: 36249, useIpv6: false, code: -98, name: EADDRINUSE, message: address already in use
The later errors are cascading startup failures:
EOFError
Rank 0 scheduler is dead
RuntimeError: Orchestrator initialization failed
The cancelled and retriggered runs overlap in their workflow timing:
- Run
34074026875: started at 2026-09-07T01:45:11Z; its GPU job ran from 01:46:20Z until cancellation at 01:53:20Z.
- Run
34074399708: created at 01:52:23Z, before the previous GPU job had completed cancellation and cleanup.
The GPU jobs from these runs, as well as the later failing run, reported the same dynamic runner name:
verl_omni_ci_554a_2026_09_07_09_45_18
This suggests that cancellation/reprovisioning may reuse a runner before all processes and listening sockets from the previous revision are fully terminated. This needs confirmation; another contributing factor may be the existing get_free_port() time-of-check/time-of-use window in vllm_omni_async_server.py, where the reservation socket is closed before AsyncOmni creates the TCPStore.
trainer.ray_master_port_range=[22000,23000] does not prevent this particular collision: the failed internal vLLM-Omni MASTER_PORT was 36249.
Expected behavior
Retriggering ready-for-ci after a PR update should safely cancel the old GPU smoke execution, fully terminate its Ray/vLLM/vLLM-Omni processes, release or destroy its dynamic runner, and start the new revision in an isolated clean environment.
The new run should not reuse stale processes or sockets from the cancelled run, and internal diffusion MASTER_PORT allocation should not race with another server.
Possible areas to investigate:
- Wait for cancelled GPU jobs and dynamic-runner cleanup to finish before provisioning or reusing a runner for the same PR.
- Ensure the dynamic runner identity is unique per workflow run and is not returned while the previous task is still shutting down.
- Add a cancellation trap that force-stops Ray and remaining vLLM/vLLM-Omni processes.
- Make internal diffusion
MASTER_PORT allocation collision-safe instead of closing the reservation socket before the TCPStore binds.
System Info
verl-omni GitHub Actions GPU smoke workflow on the dynamically provisioned
L20x8runner.Observed on PR #524:
Information
Tasks
tests/gpu_smokesuiteReproduction
ready-for-cilabel to a pull request so that the GPU smoke workflow starts.ready-for-ciagain for the updated revision.torch.distributed.TCPStorebecause the selected port is already in use.The first hard error from the affected run was:
The later errors are cascading startup failures:
The cancelled and retriggered runs overlap in their workflow timing:
34074026875: started at2026-09-07T01:45:11Z; its GPU job ran from01:46:20Zuntil cancellation at01:53:20Z.34074399708: created at01:52:23Z, before the previous GPU job had completed cancellation and cleanup.The GPU jobs from these runs, as well as the later failing run, reported the same dynamic runner name:
This suggests that cancellation/reprovisioning may reuse a runner before all processes and listening sockets from the previous revision are fully terminated. This needs confirmation; another contributing factor may be the existing
get_free_port()time-of-check/time-of-use window invllm_omni_async_server.py, where the reservation socket is closed beforeAsyncOmnicreates the TCPStore.trainer.ray_master_port_range=[22000,23000]does not prevent this particular collision: the failed internal vLLM-OmniMASTER_PORTwas36249.Expected behavior
Retriggering
ready-for-ciafter a PR update should safely cancel the old GPU smoke execution, fully terminate its Ray/vLLM/vLLM-Omni processes, release or destroy its dynamic runner, and start the new revision in an isolated clean environment.The new run should not reuse stale processes or sockets from the cancelled run, and internal diffusion
MASTER_PORTallocation should not race with another server.Possible areas to investigate:
MASTER_PORTallocation collision-safe instead of closing the reservation socket before the TCPStore binds.