Skip to content

[CI] Retriggering ready-for-ci after a PR update can cause GPU smoke port conflicts #554

Description

@Sky-Trigger

System Info

verl-omni GitHub Actions GPU smoke workflow on the dynamically provisioned L20x8 runner.

Observed on PR #524:

Information

  • The official CI scripts
  • My own modified scripts

Tasks

  • An officially supported task in the tests/gpu_smoke suite
  • My own task or dataset

Reproduction

  1. Add the ready-for-ci label to a pull request so that the GPU smoke workflow starts.
  2. While the GPU smoke job is still running, update or force-push the PR branch.
  3. Trigger ready-for-ci again for the updated revision.
  4. The previous workflow is cancelled through the workflow concurrency group, and a new dynamic GPU runner job is scheduled.
  5. During the new GPU smoke run, a diffusion worker may fail to initialize its torch.distributed.TCPStore because the selected port is already in use.

The first hard error from the affected run was:

torch.distributed.DistNetworkError: The server socket has failed to listen on any local network address. port: 36249, useIpv6: false, code: -98, name: EADDRINUSE, message: address already in use

The later errors are cascading startup failures:

EOFError
Rank 0 scheduler is dead
RuntimeError: Orchestrator initialization failed

The cancelled and retriggered runs overlap in their workflow timing:

  • Run 34074026875: started at 2026-09-07T01:45:11Z; its GPU job ran from 01:46:20Z until cancellation at 01:53:20Z.
  • Run 34074399708: created at 01:52:23Z, before the previous GPU job had completed cancellation and cleanup.

The GPU jobs from these runs, as well as the later failing run, reported the same dynamic runner name:

verl_omni_ci_554a_2026_09_07_09_45_18

This suggests that cancellation/reprovisioning may reuse a runner before all processes and listening sockets from the previous revision are fully terminated. This needs confirmation; another contributing factor may be the existing get_free_port() time-of-check/time-of-use window in vllm_omni_async_server.py, where the reservation socket is closed before AsyncOmni creates the TCPStore.

trainer.ray_master_port_range=[22000,23000] does not prevent this particular collision: the failed internal vLLM-Omni MASTER_PORT was 36249.

Expected behavior

Retriggering ready-for-ci after a PR update should safely cancel the old GPU smoke execution, fully terminate its Ray/vLLM/vLLM-Omni processes, release or destroy its dynamic runner, and start the new revision in an isolated clean environment.

The new run should not reuse stale processes or sockets from the cancelled run, and internal diffusion MASTER_PORT allocation should not race with another server.

Possible areas to investigate:

  • Wait for cancelled GPU jobs and dynamic-runner cleanup to finish before provisioning or reusing a runner for the same PR.
  • Ensure the dynamic runner identity is unique per workflow run and is not returned while the previous task is still shutting down.
  • Add a cancellation trap that force-stops Ray and remaining vLLM/vLLM-Omni processes.
  • Make internal diffusion MASTER_PORT allocation collision-safe instead of closing the reservation socket before the TCPStore binds.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't working

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions