Skip to content

[RFC] Omni V1 Trainer: Colocate Async and Separate Async Modes #320

Description

@zhtmike

Status: updated 2026-09-03, ready to implement PR 2 / PR 3 — the engine foundation this RFC was blocked on is restored: #491 step 1 (sub-RFC #492: vllm-omni pin bump + async rollout server lifecycle redesign) landed in #497, merged 2026-09-03. This RFC now tracks the remaining PR 2 (omni_colocate_async) and PR 3 (omni_separate_async) of #491's stack.
Owners: @zhtmike — colocate_async (PR 2, landing vehicle: rebase of #337) · @AndyZhou952 — separate_async (PR 3, landing vehicle: rebase of #380) — per the split recorded in this issue's comments.
Design locked at the level described here. One open design question — resume ownership vs #517 — is called out in §6.3 and must be settled before PR 3's worker hunks are finalized.


Revision note (2026-09-03) — what changed vs the 2026-07-30 body

The original body was written against the old vllm-omni pin 4444856, before #491's stack landed. It has been rewritten against the shipped state of main @ 332cb352 (#497). Substantive corrections:

  1. Engine contract replaced wholesale. The three blocking engine defects (unacked abort, frontend-flag-only pause, discarded partial tokens) are fixed upstream (vllm-omni #6084 / #6367 / #6327) and consumed by [BREAKING][vllm_omni, rollout, trainer, ci, docker, diffusion, model, tests] fix: bump vllm-omni pin to 0.28 and fix the async rollout server lifecycle #497's redesigned server. The old "Verified Compatibility" table described pre-bump semantics (e.g. bulk abort = pause_generation only) that are now wrong and is replaced by §2.
  2. Resume ownership moved server-side (deviation from sub-RFC [RFC]: Bump vLLM-Omni Pin to ded89346 and Properly Fix the Async Rollout Server Lifecycle #492, shipped in [BREAKING][vllm_omni, rollout, trainer, ci, docker, diffusion, model, tests] fix: bump vllm-omni pin to 0.28 and fix the async rollout server lifecycle #497). Every ACK-validated wake_up now resumes admission inside the server (vllm_omni_async_server.py:217), restoring upstream verl parity — AR trainers work unmodified. The trainer-visible hold now lives only on the abort path, lifted by resume_generation_replicas(). The omni_sync bridge (ray_omni_trainer.py:78-84) stays as an init-time safety net. PR 2 / PR 3 add no new per-trainer bridges; see §6.3 for the [Bug] omni admission-resume is duplicated per trainer — own it on the rollout/weight-sync path instead #517 coordination.
  3. Separate-async lifecycle description corrected. verl #7373 (2026-08-21, inside this RFC→pin window) moved hybrid↔standalone switching out of on_sample_begin/on_sample_end (now accounting-only) into on_step_begin / prepare_step / on_step_end / on_validate_begin, behind HybridRolloutSwitchConfig (enable_switch). §3.3 describes the pin-accurate hooks.
  4. parameter_sync_step bug confirmed still live — at our verl pin fefb080 (trainer_base.py:136) and on verl main as of 2026-09-03 (byte-identical). The RFC's __init__ override fix stays.
  5. OmniDetachActorWorker re-specified to the composition pattern already shipped for diffusion (DiffusionDetachActorWorker, verl_omni/workers/detach_actor_worker.py:43), matching [trainer, cfg, tests, recipe] feat: add omni separate-async V1 trainer #380's actual hunks; the old sketch routed __init__ through a single parent.
  6. adapter_name forwarding fix still open — OmniFSDPEngine.get_per_tensor_param (omni_impl.py:58) still swallows adapter_name in **kwargs; collect_lora_params already accepts it (fsdp_utils.py:422). Kept in scope (§3.7).
  7. [trainer, tests, recipe] feat: add omni colocate-async V1 trainer #337's server hunk is dead — it targets engine.abort_with_output_ids(...), which does not exist at pin ded89346; PR 2 drops it wholesale (§6.1).
  8. Environment notes refreshed: kernels stays ==0.16.0 everywhere (the ==0.14.1 plan in [RFC]: Restore Async Omni Training via the vLLM-Omni Pin Bump (Colocate & Separate Async) #491 §7 was reversed inside [BREAKING][vllm_omni, rollout, trainer, ci, docker, diffusion, model, tests] fix: bump vllm-omni pin to 0.28 and fix the async rollout server lifecycle #497 — 0.14.1 cannot select a stable-ABI FA3 build on torch 2.13 and silently falls back to native attention); VLLM_OMNI_ASYNC_OUTPUT_TIMEOUT default is 600 s (engine-side); the old drain knob is gone, replaced by VERL_OMNI_ABORT_ACK_TIMEOUT_S (default 120 s).
  9. No omni_* config stubs. The 2026-07-30 body proposed mirroring stubs; the update replaces them with a single rule — the generic keys are the read path for warmup and parameter_sync_step (§3.6) — and drops [trainer, tests, recipe] feat: add omni colocate-async V1 trainer #337's omni-key warmup override accordingly (§3.2).

1. Motivation

RFC #232 (closed; implemented across #253, #258, #312) formalized omni-model support in the verl-omni v1 trainer with OmniPPOTrainerSync (@register_trainer("omni_sync")), enabling thinker-only GSPO/DPO training with TransferQueue + ReplayBuffer. The omni trainer is algorithm-agnostic — GSPO, GRPO, etc. all work through verl's core_algos.

However, only the synchronous orchestration mode is implemented. verl's v1 trainer infrastructure supports three modes:

Mode Deployment Key characteristic Status
sync Colocated Simple: train blocks until rollout completes ✅ omni_sync
colocate_async Colocated Overlapped: rollout continues during training; aborts in-flight requests to sync weights ❌ → PR 2
separate_async Separate GPU pools Fully pipelined: dedicated rollout GPUs, zero training interruption ❌ → PR 3

The diffusion v1 trainer already has the three-mode architecture (verl_omni/trainer/diffusion/v1/trainer_separate_async.py is the precedent). Historically the omni async modes were blocked by three vllm-omni engine defects (diagnosed in #380's thread: unacked abort_async, frontend-flag-only pause_generation, bulk abort discarding partial tokens). Those are all fixed upstream and consumed by #497 — nothing engine-side blocks the trainers anymore.

2. Foundation: the shipped rollout-server contract (post-#497)

PR 2 / PR 3 build entirely on this contract. Everything in §2 is verified against the merged tree (332cb352) and pinned by CPU tests.

Pins. vllm-omni ded8934626aaad1a3e816c3a1d9d742efc012d93 (frozen 2026-08-31; ==0.28.0rc1, vLLM 0.28, torch 2.13) · verl fefb080 (2026-08-25; no verl bump required) · kernels ==0.16.0 · diffusers >=0.40.0.

2.1 Server surface (verl_omni/workers/rollout/vllm_rollout/vllm_omni_async_server.py)

Method (on vLLMOmniHttpServer) Line Semantics
abort_all_requests(reset_prefix_cache=True) -> dict :364 Abort-then-pause. Dedupes live ids from engine.request_states (:370-379); batched acked engine.abort(request_ids) bounded by VERL_OMNI_ABORT_ACK_TIMEOUT_S (default 120 s, :387-389); then pause_generation(mode="abort", wait_for_inflight_requests=False, clear_cache=reset_prefix_cache) (:392-394). Runs even with zero in-flight requests — the empty-list hazard is closed by always pausing (:391). Returns {aborted_count, request_ids} (:410). On failure: enqueue synthetic terminals for every in-flight request, then raise (:395-401). Clears frontend mm cache after pause when reset_prefix_cache=True (:403-407)
abort_request(request_id, reset_prefix_cache=True) -> dict :450 Aborts one id without pausing the engine (semantics fixed in #497 — the old RFC listed this as broken); ignores reset_prefix_cache; fail-closed (:466-469)
sleep() :233 engine.sleep(level=1) (hardcoded via _resolve_sleep_level() -> 1, :201-206 — omni default is 2, never inherit), ACK-validated, resets frontend mm cache (:225-231), invalidates LoRA cache
wake_up(tags=None) :208 engine.wake_up(tags=...) with default tags ["weights"] (:198-199), ACK-validated, then engine.resume_generation() (:217) — see §2.2, invalidates LoRA cache
release_kv_cache() / resume_kv_cache() :244 / :259 Weight-sync bracket for non-colocated modes: sleep(level=1) → wake(tags=["weights"]); then wake(tags=["kv_cache"]). Both ACK-validated, both resume admission (:256, :267)
resume_generation() :270 Bare passthrough (driver rank only)
set_global_steps(n) :220 Invalidates LoRA cache on version change
wait_for_requests_to_drain() :354 Documented no-op (TODO ... once DP is supported) — #497's commit message overstated this; the engine's _admitting gate is the sleep-time safety. verl's CheckpointEngineManager calls it before sleep; the no-op is accepted because pause_generation in abort_all_requests is the actual quiesce barrier
generate(...) :286 Delegates to the strategy layer; AR partial tokens surface through the same queues with stop_reason="aborted" (mapper at vllm_omni_strategy_base.py:225-232 — keep this frozen seam)

ACK validation (_validate_acks, :274-280) fails closed: any error-dict or non-SUCCESS OmniACK raises RuntimeError. This is load-bearing for diffusion/mixed engines (upstream re-raise covers AR EngineCore methods only) and defense-in-depth for AR-only.

Old workarounds: the 120 s natural-drain phase and VERL_OMNI_ABORT_DRAIN_TIMEOUT_S are deleted; _enqueue_abort_output (:412-448) survives as a failure-path-only synthesizer (success path consumes the engine's real cumulative-token terminals — pinned by test_abort_outputs_come_from_engine_queues_with_non_empty_tokens).

2.2 Canonical cycle and the resume-ownership map

abort (while generate is live) → pause (idle boundary + admission hold)
→ sleep(level=1) → train/idle → wake (ACK-validated; server self-resumes admission)
→ [abort path only: trainer calls resume_generation_replicas] → generate

2.3 Known limits at this pin (scope boundaries for PR 2 / PR 3)

  1. Multi-stage AR abort (thinker + talker) is broken upstream (in-code comment :383-386; needs a vllm-omni fix + pin bump). Single-stage / thinker-only rollouts — which is exactly what [trainer, tests, recipe] feat: add omni colocate-async V1 trainer #337/[trainer, cfg, tests, recipe] feat: add omni separate-async V1 trainer #380 and the registered smoke do — are correct. PR 2 / PR 3 land thinker-only first; talker-in-the-loop is out of scope (§9).
  2. Stale prefix-cache hashes after sleep/wake (#6442, still open). Mechanism addressed by #6084's EngineCore ordering; the [BREAKING][vllm_omni, rollout, trainer, ci, docker, diffusion, model, tests] fix: bump vllm-omni pin to 0.28 and fix the async rollout server lifecycle #497 merge-blocking GPU signature (first post-wake text-only-prefix request re-prefills, num_computed_tokens == 0, answers correctly, with enable_prefix_caching=True) passed. Keep that signature as a standing gate for PR 2/PR 3 GPU runs.
  3. Frontend mm-cache drift — AsyncOmni.reset_mm_cache() never finds the cache (upstream #6972, 2026-09-03); [BREAKING][vllm_omni, rollout, trainer, ci, docker, diffusion, model, tests] fix: bump vllm-omni pin to 0.28 and fix the async rollout server lifecycle #497 works around it renderer-side (:225-231, drop after #7003 lands and the pin moves).
  4. [RFC][tracking] Fail Fast — No Silent Fallbacks #388 B4 device-probe hardening was reverted from [BREAKING][vllm_omni, rollout, trainer, ci, docker, diffusion, model, tests] fix: bump vllm-omni pin to 0.28 and fix the async rollout server lifecycle #497 as out of scope; follow-up owned by [RFC][tracking] Fail Fast — No Silent Fallbacks #388.

2.4 What transfers from verl — and what must not

Practice Upstream verl (drives AsyncLLM) verl-omni server (drives AsyncOmni, post-#497)
Abort single pause_generation(wait_for_inflight_requests=False, clear_cache=…) — pause is abort + token delivery + stay-paused (vllm_async_server.py:922-925) abort-then-pause: batch abort(ids) while generate is live, then pause_generation(mode="abort", …); same reset_prefix_cache kwarg, {aborted_count, request_ids} return
Sleep level _resolve_sleep_level() may return 2 (:1116-1133) always 1 (:201-206); omni sleep() defaults to 2 and wake_up after level-2 raises NotImplementedError
Wake tags ["kv_cache", "weights"] (:1112-1114) ["weights"] (:198-199) — #337's ["kv_cache","weights"] divergence was rejected; revisit only with GPU AR-KV evidence
Wake/resume EngineCore auto-resumes; frontend holds nothing Server wake_up explicitly resumes (:217); abort-path hold lifted by trainer resume_generation_replicas
Drain wait_for_requests_to_drain (300 s, raises) no-op (:354-356); quiesce barrier = the pause inside abort_all_requests
Error handling abort errors → error dict fail-closed: enqueue terminals, then raise (#433 discipline; #433 closed alongside #497 on 2026-09-03)

The sync omni path shares the same server, so this contract covers omni_sync and both async modes alike.

3. Design

3.1 Core principle: minimal new code (unchanged)

OmniPPOTrainerSync is 20 lines including the bridge hooks (ray_omni_trainer.py:65-84); its only real override is _init_tokenizer() — wiring tokenizer/processor from OmniModelConfig instead of the HF path. The async trainers are equally thin subclasses of verl's PPOTrainerColocateAsync / PPOTrainerSeparateAsync; agent-loop lifecycle, weight sync, abort/resume, reward, advantage, loss are all inherited.

verl_omni/trainer/omni/
├── ray_omni_trainer.py                     ← existing: OmniPPOTrainerSync (+ bridge)
├── ray_omni_trainer_colocate_async.py      ← NEW (PR 2): OmniPPOTrainerColocateAsync  ~15 lines
├── ray_omni_trainer_separate_async.py      ← NEW (PR 3): OmniPPOTrainerSeparateAsync ~35 lines
└── __init__.py                             ← MOD: add exports

Registration uses verl's TRAINER_REGISTRY (trainer_base.py:1897); unregistered modes fail with ValueError: Unknown trainer '<name>' (:1918-1924); registration is import-time, triggered by main_omni.py importing verl_omni.trainer.omni (main_omni.py:39). Fresh names (omni_colocate_async, omni_separate_async) — duplicate names raise at import (:1906-1911).

3.2 PR 2 — OmniPPOTrainerColocateAsync

Subclass of PPOTrainerColocateAsync (verl trainer_colocate_async.py:26); the only override is _init_tokenizer (identical to OmniPPOTrainerSync). Registered as "omni_colocate_async". #337 also carried an on_train_begin override reading v1.omni_colocate_async.num_warmup_batches — dropped in the rebase: warmup inherits verl's hook, which reads the generic v1.colocate_async key (§3.6).

Inherited lifecycle (verified at pin fefb080):

on_init_end      → checkpoint_manager.update_weights            (:36-38)
on_train_begin   → add num_warmup_batches (v1.colocate_async)   (:40-46)
on_sample_end    → abort_replicas + sleep_replicas              (:55-59)
on_step_end      → update_weights + resume_generation_replicas  (:48-53)

abort_replicas → abort_all_requests (abort-then-pause, partial tokens preserved); sleep_replicas → level-1 sleep; update_weights ends with resume_generation (verl base.py:556) so the aborted hold is lifted before the next generate. Uses FullyAsyncLLMServerClient for abort/retry — the agent loop is unaware of interruptions; aborted generations are re-submitted with accumulated token_ids (§4, row 3). No new bridges, no server changes.

3.3 PR 3 — OmniPPOTrainerSeparateAsync

Subclass of PPOTrainerSeparateAsync (verl trainer_separate_async.py:43). Overrides _init_tokenizer (identical) plus the parameter_sync_step fix (§3.4). Registered as "omni_separate_async".

Lifecycle — corrected to the pin. verl #7373 moved the hybrid↔standalone switching out of the sample hooks; the old RFC's on_sample_begin → switch_to_rollout / on_sample_end → switch_to_trainer mapping is stale (on_sample_begin/on_sample_end at :234-242 are accounting-only). Actual hook wiring at fefb080:

on_init_end      → standalone_checkpoint_manager.update_weights
                   THEN checkpoint_manager.update_weights        (:191-194, BOTH managers)
on_train_begin   → warmup batches (v1.separate_async)            (:196-202)
on_validate_begin→ switch_to_rollout                             (:204-207)
on_step_begin    → switch_to_trainer (if lending)                (:209-224)
prepare_step     → _wait_for_sampleable_and_switch               (:244-258)
on_sample_begin/end → wait-accounting only                       (:234-242)
on_step_end      → adaptive-switch metrics; add_replicas_to_balancer
                   + clear_sticky_cache; standalone update_weights (:297-359)

where switch_to_rollout (:370-376) = update_weights + resume_generation_replicas + add_replicas_to_balancer + mode flag, and switch_to_trainer (:378-383) = remove_replicas_from_balancer + abort_replicas + sleep_replicas. Hybrid replicas join the standalone balancer via add_replicas_to_balancer (:385-390; sticky cache cleared on re-registration so load redistributes). The whole HybridRolloutSwitchConfig / enable_switch (GPU lending) machinery is available and on by config. Note the standing guards: naive checkpoint-engine backend rejected (:63-65), reward model must use a resource pool (:66-71), rollout disaggregation incompatible with enable_switch (:79-86).

Standalone vLLM-Omni pool reuses the auto-generated deploy config (from OmniRolloutPipelineBase) pointed at standalone endpoints — no adapter changes; #380 carries the working example + docs.

3.4 parameter_sync_step fix (confirmed still needed)

PPOTrainer.__init__ reads self.parameter_sync_step from config.trainer.v1.get(trainer_mode, {}).get("parameter_sync_step", 1) (verl trainer_base.py:136) — for trainer_mode="omni_separate_async" the key misses and silently yields 1, disabling Decoupled PPO (π_old stabilization via CPU save/restore, trainer_separate_async.py:157-178; train_batch_size == parameter_sync_step * ppo_mini_batch_size is asserted at :54-58 against the hardcoded v1.separate_async key). Verified unfixed on verl main as of 2026-09-03 (the read sits at trainer_base.py:138 there — two lines were added above it; the bug itself is untouched) — the override stays until an upstream fix lands:

@register_trainer("omni_separate_async")
class OmniPPOTrainerSeparateAsync(PPOTrainerSeparateAsync):
    def __init__(self, config):
        super().__init__(config)
        # Base reads v1.{trainer_mode}.parameter_sync_step (misses → 1);
        # validation reads the hardcoded v1.separate_async key. Align on the
        # validated key.
        self.parameter_sync_step = config.trainer.v1.separate_async.get(
            "parameter_sync_step", 1
        )

    def _init_tokenizer(self):
        model_config: OmniModelConfig = omega_conf_to_dataclass(
            self.config.actor_rollout_ref.model, OmniModelConfig
        )
        self.tokenizer = model_config.tokenizer
        self.processor = model_config.processor

Which key wins: v1.separate_async (the validated key) — the omni_separate_async config stub therefore omits parameter_sync_step to avoid a silently-ignored duplicate.

3.5 OmniDetachActorWorker

verl's DetachActorWorker (verl/experimental/separation/engine_workers.py:36; 169 lines, does not override init_model) inherits verl's ActorRolloutRefWorker, whose init_model dispatches via _target_/hydra-instantiate (omega_conf_to_dataclass without dataclass_type, verl/workers/engine_workers.py:538-539; the HFModelConfig annotation is a stale hint). Both paths produce a real OmniModelConfig today, but that is coincidence, not guarantee — the detach path must route through verl-omni's init_model so omni-specific behavior can never silently diverge.

The shipped precedent is composition (MRO), not __init__ re-routing — DiffusionDetachActorWorker(ActorRolloutRefWorker, DetachActorWorker) at verl_omni/workers/detach_actor_worker.py:43, with CPU-snapshot save/restore at :66-84. PR 3 adds the omni twin the same way (matching #380's actual hunks in workers/engine_workers.py / workers/omni_engine_workers.py):

class OmniDetachActorWorker(ActorRolloutRefWorker, DetachActorWorker):
    """DetachActorWorker composed with verl-omni's ActorRolloutRefWorker,
    so the omni init_model path is the single source of truth."""

verl's save/restore only touches self.actor.engine.module and self.config.actor.strategy (handler dispatch covers fsdp / fsdp2 / megatron, engine_workers.py:73-108) — both set identically by verl-omni's __init__, so all of DetachActorWorker's Decoupled-PPO logic is reused verbatim. Strategy dispatch covers "fsdp"/"fsdp2" → OmniFSDPEngine. OmniPPOTrainerSync already runs through verl-omni's init_model, so this is consistent with current omni behavior — no regression.

3.6 Config: no new stubs — the generic keys are the single read path

The generic stubs already exist via generation from verl's base (_generated_omni_trainer.yaml:1021-1045: colocate_async: {num_warmup_batches: 1}, separate_async: {num_warmup_batches: 1, parameter_sync_step: 4, hybrid_rollout: {...}}), and every consumer reads them: colocate warmup (verl trainer_colocate_async.py:43), separate-async warmup and validation (trainer_separate_async.py:53-58, :199), and — via §3.4 — parameter_sync_step. trainer_mode is a registry string, not a config key: nothing reads v1.omni_*.

PR 2 / PR 3 therefore add no omni_* config stubs. An omni_colocate_async: {num_warmup_batches: ...} stub would be silently-ignored dead config — the same misleading-duplicate hazard §3.4 avoids for parameter_sync_step. One rule covers both modes: generic keys are the single read path (this is why #337's omni-key warmup override and its stub are dropped in the rebase, §3.2). Valid trainer_mode values: omni_sync, omni_colocate_async, omni_separate_async; unknown names fail fast in get_trainer_cls (:1918-1924).

3.7 Bug fix: adapter_name forwarding (still open)

OmniFSDPEngine.get_per_tensor_param (omni_impl.py:58) accepts **kwargs but does not forward adapter_name to collect_lora_params (call site :77-81), which does accept adapter_name: str = "default" (fsdp_utils.py:422-429). Non-default LoRA adapters would broadcast wrong weights in separate_async. Fix (~3 lines) + CPU routing test ship in PR 3. (Upstream FSDPEngineWithLMHead has the same latent bug — out of scope here.)

4. Verified compatibility (post-#497 evidence)

Concern Status Evidence
Bulk abort, partial tokens ✅ abort-then-pause w/ acked batch abort, real engine terminals (token_ids == [7,8,9] survives), {"aborted_count", "request_ids"} shape — test_vllm_omni_async_server_lifecycle_on_cpu.py:115/:150; empty-list abort still pauses :133
Abort-retry resume seam ✅ strategy mapper emits "aborted" (vllm_omni_strategy_base.py:225-232); verl client resumes from prompt_ids + token_ids on stop_reason in ("aborted","abort") (verl llm_server.py:404-455) — frozen seam, do not touch
Sleep/wake offload ✅ level=1 + keyword tags= + ACK fail-closed + LoRA-cache invalidation, uniform across HYBRID/COLOCATED — lifecycle tests :171/:189/:299/:323
Wake-resume parity ✅ server resumes admission on every successful wake (:217, tests :536/:546/:557); AR-only engines need no trainer resume after sleep/wake
Weight update (NCCL+ZMQ) ✅ update_weights ordering ends resume_generation (verl base.py:506-558, resume at :556); colocate on_step_end resume is idempotent (:53)
Standalone HTTP server + balancer ✅ vLLMOmniHttpServer as ray.remote actor; GlobalRequestLoadBalancer.add_servers/remove_servers + sticky-cache invalidation (verl llm_server.py:92-95/:123-149/:159-178)
Multimodal TQ serialization ✅ AgentLoopWorker._compute_multi_modal_inputs (verl agent_loop.py:838-894) → TransferQueue postprocess (agent_loop_tq.py:150-228)
GSPO rollout correction ✅ pipelines reuse saved input_ids tensors — no re-tokenization (unchanged since #232 series)
OmniDetachActorWorker + OmniFSDPEngine ✅ (by construction) composition per DiffusionDetachActorWorker (detach_actor_worker.py:43); DetachActorWorker save/restore surface verified at pin (engine_workers.py:123-161)
CheckpointEngineManager LoRA sync (default) ✅ OmniCheckpointEngineManager (checkpoint_engine.py:19) inherits; both sides land on adapter_name="default" — non-default is what §3.7 fixes
omni_sync through the bump ✅ bridge CPU tests (test_omni_sync_resume_bridge_on_cpu.py:75/:82/:89/:109/:128); AVQA GSPO 240-step GPU run green, +10-12% throughput (#497)

5. Risks & mitigations

Risk Severity Mitigation
Resume-ownership overlap with #517 (move resume to rollout/weight-sync path, delete per-trainer bridges; cites the duplicated pattern in diffusion trainer_separate_async.py:315/:324) Medium PR 2/PR 3 add no new bridges (wake self-resume + update_weights tail already cover them). #380's worker hunks are the overlap zone — settle #517's placement before PR 3's final review; PR 2 is unaffected
Lifecycle RPC starvation under Ray max_concurrency (control RPCs queued behind generates) Medium watch step time in the PR 2 GPU gate; dedicated control concurrency group is the #377 remainder (salvage #337's _CONTROL_CONCURRENCY_GROUP there, not here)
Multi-stage AR abort broken upstream (thinker+talker) Medium scope: thinker-only/single-stage for both PRs (§2.3.1); tracked for a future pin bump
Stale prefix hashes after sleep/wake (#6442, open) Medium mechanism fixed; keep the dedicated enable_prefix_caching=True GPU signature as a standing gate; recipes set prefix caching explicitly (AR default silently True)
Frontend mm-cache drift (#6972) Low renderer-side clear shipped (:225-231); drop after upstream #7003 + pin bump
parameter_sync_step silently 1 Low §3.4 override + CPU test asserting the read path; revisit when verl fixes trainer_base.py:136 upstream
adapter_name forwarding Low §3.7 fix + CPU routing test
verl-main deltas not at pin: #7511 (abort submission gate, DP>1), #7115 (balancer moved to verl/workers/rollout/router.py), #7227 (delta-sharded sync) Low pin stands (#492: no verl bump); note for the next bump — anything importing GlobalRequestLoadBalancer from llm_server adapts then
Same-file rebase friction Low #522 (BREAKING legacy Qwen3-Omni removal), #403 series (#480/#481 must rebase onto #497 first), #483/#428/#354/#464/#485/#470/#375/#486 already notified per #491 §8

6. Implementation plan & coordination

6.1 PR 2 — omni_colocate_async (rebase #337; owner @zhtmike)

#337 (draft, last push 2026-08-14, pre-bump) splits cleanly:

Gate (per #491 §6): MMK12-style colocate run — abort+sleep boundary observed, step time without drain compensation (the drain knob no longer exists; 120 s/step stalls would indicate a real bug), mm_processor_cache_gb=0.

6.2 PR 3 — omni_separate_async (rebase #380; owner @AndyZhou952)

#380 (draft, updated 2026-09-01, +907/−15 over 14 files) needs no server changes — its three historical engine blockers are fixed. Rebase keeps: trainer (+ export), workers/engine_workers.py / workers/omni_engine_workers.py / engine/fsdp/omni_impl.py (detach wiring + §3.7), docs (separate_async_omni.md), example + smoke scripts, 4 CPU tests. Two open threads from its discussion: convergence parity and "actor-updating takes extensive time" — the latter was the old frontend-only-pause symptom, expected to disappear against the new pin (verify in the gate).

Gate (per #491 §6 / #399): hard-abort/weight-sync gate (partial-token continuation, no stale/late outputs) + a convergence run with correct parameter_sync_step (Decoupled PPO engaged, not silently 1).

6.3 #517 coordination (open design question)

#517 (2026-09-03, @knlnguyen1802) proposes moving resume ownership onto the rollout/weight-sync path and deleting the per-trainer bridges (including the duplicated pattern in diffusion trainer_separate_async.py:315/:324 and, eventually, the omni_sync safety net). Neither PR 2 nor PR 3 should grow bridge code that #517 would immediately delete: colocate needs none (wake self-resume + update_weights tail suffice), separate inherits verl's switch_to_rollout resume. Sequence: land PR 2/PR 3 bridge-free; #517 consolidates server-side; then delete the omni_sync safety net.

6.4 Sequencing (mirrors #491 §4)

Both PRs depend only on landed work (1A/1B ✅, 1C ✅) and can proceed in parallel after rebase. Post-stack follow-ups (not here): #452 (unblocked — upstream #6645 merged 2026-09-01; open non-draft after), #377 remainder (concurrency group), #391, #388 B4, #6442 close-out.

Files changed:

File Change PR
verl_omni/trainer/omni/ray_omni_trainer_colocate_async.py New (~15) 2
verl_omni/trainer/omni/ray_omni_trainer_separate_async.py New (~35, incl. §3.4) 3
verl_omni/trainer/omni/__init__.py Exports 2/3
— config files No changes — no omni_* stubs; the generic keys are the read path (§3.6) —
verl_omni/workers/engine_workers.py / omni_engine_workers.py OmniDetachActorWorker (§3.5) 3
verl_omni/workers/engine/fsdp/omni_impl.py adapter_name forwarding (~3) 3
examples/gspo_trainer/qwen3_omni/run_*_colocate_async.sh MMK12-style example 2
examples/ separate-async standalone+hybrid scripts, docs/ Example + separate_async_omni.md 3
tests/trainer/omni/test_omni_*_async_on_cpu.py Registration/tokenizer/warmup/param-sync/adapter/detach tests 2/3

7. Verification

CPU (pattern: tests/**/*_on_cpu.py, --asyncio-mode=auto):

  1. Registration: get_trainer_cls("omni_colocate_async"/"omni_separate_async") resolve; unknown mode raises ValueError.
  2. _init_tokenizer wiring (processor + tokenizer from OmniModelConfig) for both trainers.
  3. Colocate warmup flows from the generic v1.colocate_async.num_warmup_batches (no override; an omni_colocate_async key, if present, is ignored).
  4. parameter_sync_step read from the validated v1.separate_async key (mock config where the keys differ; assert not silently 1).
  5. adapter_name routing: get_per_tensor_param(adapter_name="old") reaches collect_lora_params.
  6. OmniDetachActorWorker: routes through omni init_model (isinstance(worker.actor.engine, OmniFSDPEngine), model_config is OmniModelConfig); save_model_to_cpu/restore_model_from_cpu round-trip.
  7. Config: both trainer_mode values resolve end-to-end with only the generic stubs present (no omni_* keys required).

GPU gates (per #491 §6): PR 2 — MMK12-style colocate run (§6.1). PR 3 — hard-abort/weight-sync gate + convergence run (§6.2). Standing regressions: #6442 prefix signature with explicit enable_prefix_caching=True; CUDA-free server PID after wake. Scope: thinker-only/single-stage AR (§2.3.1).

8. Future work

9. Out of scope

10. Tracking checklist


Related: #491 (umbrella — PR 2/PR 3 rows), #497 (landed engine foundation), #492 (sub-RFC, closed), #337 / #380 (landing vehicles), #517 (resume ownership), #399 / #413 (finalize against the bumped pin), #6442 / #6972 / #7003 (upstream), #445 (patch ledger).

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Labels

enhancementNew feature or request

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions