You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Status: updated 2026-09-03, ready to implement PR 2 / PR 3 — the engine foundation this RFC was blocked on is restored: #491 step 1 (sub-RFC #492: vllm-omni pin bump + async rollout server lifecycle redesign) landed in #497, merged 2026-09-03. This RFC now tracks the remaining PR 2 (omni_colocate_async) and PR 3 (omni_separate_async) of #491's stack.
Owners: @zhtmike — colocate_async (PR 2, landing vehicle: rebase of #337) · @AndyZhou952 — separate_async (PR 3, landing vehicle: rebase of #380) — per the split recorded in this issue's comments.
Design locked at the level described here. One open design question — resume ownership vs #517 — is called out in §6.3 and must be settled before PR 3's worker hunks are finalized.
Revision note (2026-09-03) — what changed vs the 2026-07-30 body
The original body was written against the old vllm-omni pin 4444856, before #491's stack landed. It has been rewritten against the shipped state of main @ 332cb352 (#497). Substantive corrections:
Separate-async lifecycle description corrected. verl #7373 (2026-08-21, inside this RFC→pin window) moved hybrid↔standalone switching out of on_sample_begin/on_sample_end (now accounting-only) into on_step_begin / prepare_step / on_step_end / on_validate_begin, behind HybridRolloutSwitchConfig (enable_switch). §3.3 describes the pin-accurate hooks.
parameter_sync_step bug confirmed still live — at our verl pin fefb080 (trainer_base.py:136) and on verl main as of 2026-09-03 (byte-identical). The RFC's __init__ override fix stays.
OmniDetachActorWorker re-specified to the composition pattern already shipped for diffusion (DiffusionDetachActorWorker, verl_omni/workers/detach_actor_worker.py:43), matching [trainer, cfg, tests, recipe] feat: add omni separate-async V1 trainer #380's actual hunks; the old sketch routed __init__ through a single parent.
adapter_name forwarding fix still open — OmniFSDPEngine.get_per_tensor_param (omni_impl.py:58) still swallows adapter_name in **kwargs; collect_lora_params already accepts it (fsdp_utils.py:422). Kept in scope (§3.7).
No omni_* config stubs. The 2026-07-30 body proposed mirroring stubs; the update replaces them with a single rule — the generic keys are the read path for warmup and parameter_sync_step (§3.6) — and drops [trainer, tests, recipe] feat: add omni colocate-async V1 trainer #337's omni-key warmup override accordingly (§3.2).
1. Motivation
RFC #232 (closed; implemented across #253, #258, #312) formalized omni-model support in the verl-omni v1 trainer with OmniPPOTrainerSync (@register_trainer("omni_sync")), enabling thinker-only GSPO/DPO training with TransferQueue + ReplayBuffer. The omni trainer is algorithm-agnostic — GSPO, GRPO, etc. all work through verl's core_algos.
However, only the synchronous orchestration mode is implemented. verl's v1 trainer infrastructure supports three modes:
Mode
Deployment
Key characteristic
Status
sync
Colocated
Simple: train blocks until rollout completes
✅ omni_sync
colocate_async
Colocated
Overlapped: rollout continues during training; aborts in-flight requests to sync weights
❌ → PR 2
separate_async
Separate GPU pools
Fully pipelined: dedicated rollout GPUs, zero training interruption
❌ → PR 3
The diffusion v1 trainer already has the three-mode architecture (verl_omni/trainer/diffusion/v1/trainer_separate_async.py is the precedent). Historically the omni async modes were blocked by three vllm-omni engine defects (diagnosed in #380's thread: unacked abort_async, frontend-flag-only pause_generation, bulk abort discarding partial tokens). Those are all fixed upstream and consumed by #497 — nothing engine-side blocks the trainers anymore.
2. Foundation: the shipped rollout-server contract (post-#497)
PR 2 / PR 3 build entirely on this contract. Everything in §2 is verified against the merged tree (332cb352) and pinned by CPU tests.
Abort-then-pause. Dedupes live ids from engine.request_states (:370-379); batched acked engine.abort(request_ids) bounded by VERL_OMNI_ABORT_ACK_TIMEOUT_S (default 120 s, :387-389); then pause_generation(mode="abort", wait_for_inflight_requests=False, clear_cache=reset_prefix_cache) (:392-394). Runs even with zero in-flight requests — the empty-list hazard is closed by always pausing (:391). Returns {aborted_count, request_ids} (:410). On failure: enqueue synthetic terminals for every in-flight request, then raise (:395-401). Clears frontend mm cache after pause when reset_prefix_cache=True (:403-407)
Aborts one id without pausing the engine (semantics fixed in #497 — the old RFC listed this as broken); ignores reset_prefix_cache; fail-closed (:466-469)
sleep()
:233
engine.sleep(level=1) (hardcoded via _resolve_sleep_level() -> 1, :201-206 — omni default is 2, never inherit), ACK-validated, resets frontend mm cache (:225-231), invalidates LoRA cache
wake_up(tags=None)
:208
engine.wake_up(tags=...) with default tags ["weights"] (:198-199), ACK-validated, then engine.resume_generation() (:217) — see §2.2, invalidates LoRA cache
release_kv_cache() / resume_kv_cache()
:244 / :259
Weight-sync bracket for non-colocated modes: sleep(level=1) → wake(tags=["weights"]); then wake(tags=["kv_cache"]). Both ACK-validated, both resume admission (:256, :267)
resume_generation()
:270
Bare passthrough (driver rank only)
set_global_steps(n)
:220
Invalidates LoRA cache on version change
wait_for_requests_to_drain()
:354
Documented no-op (TODO ... once DP is supported) — #497's commit message overstated this; the engine's _admitting gate is the sleep-time safety. verl's CheckpointEngineManager calls it before sleep; the no-op is accepted because pause_generation in abort_all_requests is the actual quiesce barrier
generate(...)
:286
Delegates to the strategy layer; AR partial tokens surface through the same queues with stop_reason="aborted" (mapper at vllm_omni_strategy_base.py:225-232 — keep this frozen seam)
ACK validation (_validate_acks, :274-280) fails closed: any error-dict or non-SUCCESSOmniACK raises RuntimeError. This is load-bearing for diffusion/mixed engines (upstream re-raise covers AR EngineCore methods only) and defense-in-depth for AR-only.
Old workarounds: the 120 s natural-drain phase and VERL_OMNI_ABORT_DRAIN_TIMEOUT_S are deleted; _enqueue_abort_output (:412-448) survives as a failure-path-only synthesizer (success path consumes the engine's real cumulative-token terminals — pinned by test_abort_outputs_come_from_engine_queues_with_non_empty_tokens).
The abort path still holds until resume_generation_replicas() — verl's CheckpointEngineManager.update_weights already ends with resume_generation (verl checkpoint_engine/base.py:506-558, resume at :556), and PPOTrainerColocateAsync.on_step_end calls it again (:53) idempotently.
single pause_generation(wait_for_inflight_requests=False, clear_cache=…) — pause is abort + token delivery + stay-paused (vllm_async_server.py:922-925)
abort-then-pause: batch abort(ids) while generate is live, then pause_generation(mode="abort", …); same reset_prefix_cache kwarg, {aborted_count, request_ids} return
Sleep level
_resolve_sleep_level() may return 2 (:1116-1133)
always 1 (:201-206); omni sleep() defaults to 2 and wake_up after level-2 raises NotImplementedError
Wake tags
["kv_cache", "weights"] (:1112-1114)
["weights"] (:198-199) — #337's ["kv_cache","weights"] divergence was rejected; revisit only with GPU AR-KV evidence
Wake/resume
EngineCore auto-resumes; frontend holds nothing
Server wake_upexplicitly resumes (:217); abort-path hold lifted by trainer resume_generation_replicas
Drain
wait_for_requests_to_drain (300 s, raises)
no-op (:354-356); quiesce barrier = the pause inside abort_all_requests
Error handling
abort errors → error dict
fail-closed: enqueue terminals, then raise (#433 discipline; #433 closed alongside #497 on 2026-09-03)
The sync omni path shares the same server, so this contract covers omni_sync and both async modes alike.
3. Design
3.1 Core principle: minimal new code (unchanged)
OmniPPOTrainerSync is 20 lines including the bridge hooks (ray_omni_trainer.py:65-84); its only real override is _init_tokenizer() — wiring tokenizer/processor from OmniModelConfig instead of the HF path. The async trainers are equally thin subclasses of verl's PPOTrainerColocateAsync / PPOTrainerSeparateAsync; agent-loop lifecycle, weight sync, abort/resume, reward, advantage, loss are all inherited.
Registration uses verl's TRAINER_REGISTRY (trainer_base.py:1897); unregistered modes fail with ValueError: Unknown trainer '<name>' (:1918-1924); registration is import-time, triggered by main_omni.py importing verl_omni.trainer.omni (main_omni.py:39). Fresh names (omni_colocate_async, omni_separate_async) — duplicate names raise at import (:1906-1911).
3.2 PR 2 — OmniPPOTrainerColocateAsync
Subclass of PPOTrainerColocateAsync (verl trainer_colocate_async.py:26); the only override is _init_tokenizer (identical to OmniPPOTrainerSync). Registered as "omni_colocate_async". #337 also carried an on_train_begin override reading v1.omni_colocate_async.num_warmup_batches — dropped in the rebase: warmup inherits verl's hook, which reads the generic v1.colocate_async key (§3.6).
abort_replicas → abort_all_requests (abort-then-pause, partial tokens preserved); sleep_replicas → level-1 sleep; update_weights ends with resume_generation (verl base.py:556) so the aborted hold is lifted before the next generate. Uses FullyAsyncLLMServerClient for abort/retry — the agent loop is unaware of interruptions; aborted generations are re-submitted with accumulated token_ids (§4, row 3). No new bridges, no server changes.
3.3 PR 3 — OmniPPOTrainerSeparateAsync
Subclass of PPOTrainerSeparateAsync (verl trainer_separate_async.py:43). Overrides _init_tokenizer (identical) plus the parameter_sync_step fix (§3.4). Registered as "omni_separate_async".
Lifecycle — corrected to the pin. verl #7373 moved the hybrid↔standalone switching out of the sample hooks; the old RFC's on_sample_begin → switch_to_rollout / on_sample_end → switch_to_trainer mapping is stale (on_sample_begin/on_sample_end at :234-242 are accounting-only). Actual hook wiring at fefb080:
where switch_to_rollout (:370-376) = update_weights + resume_generation_replicas + add_replicas_to_balancer + mode flag, and switch_to_trainer (:378-383) = remove_replicas_from_balancer + abort_replicas + sleep_replicas. Hybrid replicas join the standalone balancer via add_replicas_to_balancer (:385-390; sticky cache cleared on re-registration so load redistributes). The whole HybridRolloutSwitchConfig / enable_switch (GPU lending) machinery is available and on by config. Note the standing guards: naive checkpoint-engine backend rejected (:63-65), reward model must use a resource pool (:66-71), rollout disaggregation incompatible with enable_switch (:79-86).
Standalone vLLM-Omni pool reuses the auto-generated deploy config (from OmniRolloutPipelineBase) pointed at standalone endpoints — no adapter changes; #380 carries the working example + docs.
3.4 parameter_sync_step fix (confirmed still needed)
PPOTrainer.__init__ reads self.parameter_sync_step from config.trainer.v1.get(trainer_mode, {}).get("parameter_sync_step", 1) (verl trainer_base.py:136) — for trainer_mode="omni_separate_async" the key misses and silently yields 1, disabling Decoupled PPO (π_old stabilization via CPU save/restore, trainer_separate_async.py:157-178; train_batch_size == parameter_sync_step * ppo_mini_batch_size is asserted at :54-58 against the hardcoded v1.separate_async key). Verified unfixed on verl main as of 2026-09-03 (the read sits at trainer_base.py:138 there — two lines were added above it; the bug itself is untouched) — the override stays until an upstream fix lands:
Which key wins:v1.separate_async (the validated key) — the omni_separate_async config stub therefore omitsparameter_sync_step to avoid a silently-ignored duplicate.
3.5 OmniDetachActorWorker
verl's DetachActorWorker (verl/experimental/separation/engine_workers.py:36; 169 lines, does not override init_model) inherits verl's ActorRolloutRefWorker, whose init_model dispatches via _target_/hydra-instantiate (omega_conf_to_dataclass without dataclass_type, verl/workers/engine_workers.py:538-539; the HFModelConfig annotation is a stale hint). Both paths produce a real OmniModelConfig today, but that is coincidence, not guarantee — the detach path must route through verl-omni's init_model so omni-specific behavior can never silently diverge.
The shipped precedent is composition (MRO), not __init__ re-routing — DiffusionDetachActorWorker(ActorRolloutRefWorker, DetachActorWorker) at verl_omni/workers/detach_actor_worker.py:43, with CPU-snapshot save/restore at :66-84. PR 3 adds the omni twin the same way (matching #380's actual hunks in workers/engine_workers.py / workers/omni_engine_workers.py):
classOmniDetachActorWorker(ActorRolloutRefWorker, DetachActorWorker):
"""DetachActorWorker composed with verl-omni's ActorRolloutRefWorker, so the omni init_model path is the single source of truth."""
verl's save/restore only touches self.actor.engine.module and self.config.actor.strategy (handler dispatch covers fsdp / fsdp2 / megatron, engine_workers.py:73-108) — both set identically by verl-omni's __init__, so all of DetachActorWorker's Decoupled-PPO logic is reused verbatim. Strategy dispatch covers "fsdp"/"fsdp2" → OmniFSDPEngine. OmniPPOTrainerSync already runs through verl-omni's init_model, so this is consistent with current omni behavior — no regression.
3.6 Config: no new stubs — the generic keys are the single read path
The generic stubs already exist via generation from verl's base (_generated_omni_trainer.yaml:1021-1045: colocate_async: {num_warmup_batches: 1}, separate_async: {num_warmup_batches: 1, parameter_sync_step: 4, hybrid_rollout: {...}}), and every consumer reads them: colocate warmup (verl trainer_colocate_async.py:43), separate-async warmup and validation (trainer_separate_async.py:53-58, :199), and — via §3.4 — parameter_sync_step. trainer_mode is a registry string, not a config key: nothing reads v1.omni_*.
PR 2 / PR 3 therefore add no omni_* config stubs. An omni_colocate_async: {num_warmup_batches: ...} stub would be silently-ignored dead config — the same misleading-duplicate hazard §3.4 avoids for parameter_sync_step. One rule covers both modes: generic keys are the single read path (this is why #337's omni-key warmup override and its stub are dropped in the rebase, §3.2). Valid trainer_mode values: omni_sync, omni_colocate_async, omni_separate_async; unknown names fail fast in get_trainer_cls (:1918-1924).
3.7 Bug fix: adapter_name forwarding (still open)
OmniFSDPEngine.get_per_tensor_param (omni_impl.py:58) accepts **kwargs but does not forward adapter_name to collect_lora_params (call site :77-81), which does accept adapter_name: str = "default" (fsdp_utils.py:422-429). Non-default LoRA adapters would broadcast wrong weights in separate_async. Fix (~3 lines) + CPU routing test ship in PR 3. (Upstream FSDPEngineWithLMHead has the same latent bug — out of scope here.)
strategy mapper emits "aborted" (vllm_omni_strategy_base.py:225-232); verl client resumes from prompt_ids + token_ids on stop_reason in ("aborted","abort") (verl llm_server.py:404-455) — frozen seam, do not touch
pipelines reuse saved input_ids tensors — no re-tokenization (unchanged since #232 series)
OmniDetachActorWorker + OmniFSDPEngine
✅ (by construction)
composition per DiffusionDetachActorWorker (detach_actor_worker.py:43); DetachActorWorker save/restore surface verified at pin (engine_workers.py:123-161)
CheckpointEngineManager LoRA sync (default)
✅
OmniCheckpointEngineManager (checkpoint_engine.py:19) inherits; both sides land on adapter_name="default" — non-default is what §3.7 fixes
omni_sync through the bump
✅
bridge CPU tests (test_omni_sync_resume_bridge_on_cpu.py:75/:82/:89/:109/:128); AVQA GSPO 240-step GPU run green, +10-12% throughput (#497)
5. Risks & mitigations
Risk
Severity
Mitigation
Resume-ownership overlap with #517 (move resume to rollout/weight-sync path, delete per-trainer bridges; cites the duplicated pattern in diffusion trainer_separate_async.py:315/:324)
Medium
PR 2/PR 3 add no new bridges (wake self-resume + update_weights tail already cover them). #380's worker hunks are the overlap zone — settle #517's placement before PR 3's final review; PR 2 is unaffected
Lifecycle RPC starvation under Ray max_concurrency (control RPCs queued behind generates)
Medium
watch step time in the PR 2 GPU gate; dedicated control concurrency group is the #377 remainder (salvage #337's _CONTROL_CONCURRENCY_GROUP there, not here)
Multi-stage AR abort broken upstream (thinker+talker)
Medium
scope: thinker-only/single-stage for both PRs (§2.3.1); tracked for a future pin bump
Stale prefix hashes after sleep/wake (#6442, open)
Medium
mechanism fixed; keep the dedicated enable_prefix_caching=True GPU signature as a standing gate; recipes set prefix caching explicitly (AR default silently True)
Frontend mm-cache drift (#6972)
Low
renderer-side clear shipped (:225-231); drop after upstream #7003 + pin bump
parameter_sync_step silently 1
Low
§3.4 override + CPU test asserting the read path; revisit when verl fixes trainer_base.py:136 upstream
adapter_name forwarding
Low
§3.7 fix + CPU routing test
verl-main deltas not at pin: #7511 (abort submission gate, DP>1), #7115 (balancer moved to verl/workers/rollout/router.py), #7227 (delta-sharded sync)
Low
pin stands (#492: no verl bump); note for the next bump — anything importing GlobalRequestLoadBalancer from llm_server adapts then
Gate (per #491 §6): MMK12-style colocate run — abort+sleep boundary observed, step time without drain compensation (the drain knob no longer exists; 120 s/step stalls would indicate a real bug), mm_processor_cache_gb=0.
#380 (draft, updated 2026-09-01, +907/−15 over 14 files) needs no server changes — its three historical engine blockers are fixed. Rebase keeps: trainer (+ export), workers/engine_workers.py / workers/omni_engine_workers.py / engine/fsdp/omni_impl.py (detach wiring + §3.7), docs (separate_async_omni.md), example + smoke scripts, 4 CPU tests. Two open threads from its discussion: convergence parity and "actor-updating takes extensive time" — the latter was the old frontend-only-pause symptom, expected to disappear against the new pin (verify in the gate).
Gate (per #491 §6 / #399): hard-abort/weight-sync gate (partial-token continuation, no stale/late outputs) + a convergence run with correct parameter_sync_step (Decoupled PPO engaged, not silently 1).
Both PRs depend only on landed work (1A/1B ✅, 1C ✅) and can proceed in parallel after rebase. Post-stack follow-ups (not here): #452 (unblocked — upstream #6645 merged 2026-09-01; open non-draft after), #377 remainder (concurrency group), #391, #388 B4, #6442 close-out.
OmniDetachActorWorker: routes through omni init_model (isinstance(worker.actor.engine, OmniFSDPEngine), model_config is OmniModelConfig); save_model_to_cpu/restore_model_from_cpu round-trip.
Config: both trainer_mode values resolve end-to-end with only the generic stubs present (no omni_* keys required).
GPU gates (per #491 §6): PR 2 — MMK12-style colocate run (§6.1). PR 3 — hard-abort/weight-sync gate + convergence run (§6.2). Standing regressions: #6442 prefix signature with explicit enable_prefix_caching=True; CUDA-free server PID after wake. Scope: thinker-only/single-stage AR (§2.3.1).
Status: updated 2026-09-03, ready to implement PR 2 / PR 3 — the engine foundation this RFC was blocked on is restored: #491 step 1 (sub-RFC #492: vllm-omni pin bump + async rollout server lifecycle redesign) landed in #497, merged 2026-09-03. This RFC now tracks the remaining PR 2 (
omni_colocate_async) and PR 3 (omni_separate_async) of #491's stack.Owners: @zhtmike — colocate_async (PR 2, landing vehicle: rebase of #337) · @AndyZhou952 — separate_async (PR 3, landing vehicle: rebase of #380) — per the split recorded in this issue's comments.
Design locked at the level described here. One open design question — resume ownership vs #517 — is called out in §6.3 and must be settled before PR 3's worker hunks are finalized.
Revision note (2026-09-03) — what changed vs the 2026-07-30 body
The original body was written against the old vllm-omni pin
4444856, before #491's stack landed. It has been rewritten against the shipped state ofmain@332cb352(#497). Substantive corrections:pause_generationonly) that are now wrong and is replaced by §2.ded89346and Properly Fix the Async Rollout Server Lifecycle #492, shipped in [BREAKING][vllm_omni, rollout, trainer, ci, docker, diffusion, model, tests] fix: bump vllm-omni pin to 0.28 and fix the async rollout server lifecycle #497). Every ACK-validatedwake_upnow resumes admission inside the server (vllm_omni_async_server.py:217), restoring upstream verl parity — AR trainers work unmodified. The trainer-visible hold now lives only on the abort path, lifted byresume_generation_replicas(). Theomni_syncbridge (ray_omni_trainer.py:78-84) stays as an init-time safety net. PR 2 / PR 3 add no new per-trainer bridges; see §6.3 for the [Bug] omni admission-resume is duplicated per trainer — own it on the rollout/weight-sync path instead #517 coordination.on_sample_begin/on_sample_end(now accounting-only) intoon_step_begin/prepare_step/on_step_end/on_validate_begin, behindHybridRolloutSwitchConfig(enable_switch). §3.3 describes the pin-accurate hooks.parameter_sync_stepbug confirmed still live — at our verl pinfefb080(trainer_base.py:136) and on verl main as of 2026-09-03 (byte-identical). The RFC's__init__override fix stays.OmniDetachActorWorkerre-specified to the composition pattern already shipped for diffusion (DiffusionDetachActorWorker,verl_omni/workers/detach_actor_worker.py:43), matching [trainer, cfg, tests, recipe] feat: add omni separate-async V1 trainer #380's actual hunks; the old sketch routed__init__through a single parent.adapter_nameforwarding fix still open —OmniFSDPEngine.get_per_tensor_param(omni_impl.py:58) still swallowsadapter_namein**kwargs;collect_lora_paramsalready accepts it (fsdp_utils.py:422). Kept in scope (§3.7).engine.abort_with_output_ids(...), which does not exist at pinded89346; PR 2 drops it wholesale (§6.1).==0.16.0everywhere (the==0.14.1plan in [RFC]: Restore Async Omni Training via the vLLM-Omni Pin Bump (Colocate & Separate Async) #491 §7 was reversed inside [BREAKING][vllm_omni, rollout, trainer, ci, docker, diffusion, model, tests] fix: bump vllm-omni pin to 0.28 and fix the async rollout server lifecycle #497 — 0.14.1 cannot select a stable-ABI FA3 build on torch 2.13 and silently falls back to native attention);VLLM_OMNI_ASYNC_OUTPUT_TIMEOUTdefault is 600 s (engine-side); the old drain knob is gone, replaced byVERL_OMNI_ABORT_ACK_TIMEOUT_S(default 120 s).omni_*config stubs. The 2026-07-30 body proposed mirroring stubs; the update replaces them with a single rule — the generic keys are the read path for warmup andparameter_sync_step(§3.6) — and drops [trainer, tests, recipe] feat: add omni colocate-async V1 trainer #337's omni-key warmup override accordingly (§3.2).1. Motivation
RFC #232 (closed; implemented across #253, #258, #312) formalized omni-model support in the verl-omni v1 trainer with
OmniPPOTrainerSync(@register_trainer("omni_sync")), enabling thinker-only GSPO/DPO training with TransferQueue + ReplayBuffer. The omni trainer is algorithm-agnostic — GSPO, GRPO, etc. all work through verl'score_algos.However, only the synchronous orchestration mode is implemented. verl's v1 trainer infrastructure supports three modes:
syncomni_synccolocate_asyncseparate_asyncThe diffusion v1 trainer already has the three-mode architecture (
verl_omni/trainer/diffusion/v1/trainer_separate_async.pyis the precedent). Historically the omni async modes were blocked by three vllm-omni engine defects (diagnosed in #380's thread: unackedabort_async, frontend-flag-onlypause_generation, bulk abort discarding partial tokens). Those are all fixed upstream and consumed by #497 — nothing engine-side blocks the trainers anymore.2. Foundation: the shipped rollout-server contract (post-#497)
PR 2 / PR 3 build entirely on this contract. Everything in §2 is verified against the merged tree (
332cb352) and pinned by CPU tests.Pins. vllm-omni
ded8934626aaad1a3e816c3a1d9d742efc012d93(frozen 2026-08-31;==0.28.0rc1, vLLM 0.28, torch 2.13) · verlfefb080(2026-08-25; no verl bump required) · kernels==0.16.0· diffusers>=0.40.0.2.1 Server surface (
verl_omni/workers/rollout/vllm_rollout/vllm_omni_async_server.py)vLLMOmniHttpServer)abort_all_requests(reset_prefix_cache=True) -> dict:364engine.request_states(:370-379); batched ackedengine.abort(request_ids)bounded byVERL_OMNI_ABORT_ACK_TIMEOUT_S(default 120 s, :387-389); thenpause_generation(mode="abort", wait_for_inflight_requests=False, clear_cache=reset_prefix_cache)(:392-394). Runs even with zero in-flight requests — the empty-list hazard is closed by always pausing (:391). Returns{aborted_count, request_ids}(:410). On failure: enqueue synthetic terminals for every in-flight request, then raise (:395-401). Clears frontend mm cache after pause whenreset_prefix_cache=True(:403-407)abort_request(request_id, reset_prefix_cache=True) -> dict:450reset_prefix_cache; fail-closed (:466-469)sleep():233engine.sleep(level=1)(hardcoded via_resolve_sleep_level() -> 1, :201-206 — omni default is 2, never inherit), ACK-validated, resets frontend mm cache (:225-231), invalidates LoRA cachewake_up(tags=None):208engine.wake_up(tags=...)with default tags["weights"](:198-199), ACK-validated, thenengine.resume_generation()(:217) — see §2.2, invalidates LoRA cacherelease_kv_cache()/resume_kv_cache():244/:259["weights"]); then wake(tags=["kv_cache"]). Both ACK-validated, both resume admission (:256, :267)resume_generation():270set_global_steps(n):220wait_for_requests_to_drain():354TODO ... once DP is supported) — #497's commit message overstated this; the engine's_admittinggate is the sleep-time safety. verl'sCheckpointEngineManagercalls it before sleep; the no-op is accepted becausepause_generationinabort_all_requestsis the actual quiesce barriergenerate(...):286stop_reason="aborted"(mapper atvllm_omni_strategy_base.py:225-232— keep this frozen seam)ACK validation (
_validate_acks, :274-280) fails closed: any error-dict or non-SUCCESSOmniACKraisesRuntimeError. This is load-bearing for diffusion/mixed engines (upstream re-raise covers AR EngineCore methods only) and defense-in-depth for AR-only.Old workarounds: the 120 s natural-drain phase and
VERL_OMNI_ABORT_DRAIN_TIMEOUT_Sare deleted;_enqueue_abort_output(:412-448) survives as a failure-path-only synthesizer (success path consumes the engine's real cumulative-token terminals — pinned bytest_abort_outputs_come_from_engine_queues_with_non_empty_tokens).2.2 Canonical cycle and the resume-ownership map
:217, also:256/:267): restores upstream verl's contract (vLLMAsyncLLMholds nothing across sleep/wake), so every AR trainer works unmodified — no per-mode resume hooks are needed after sleep/wake. This is a documented deviation from sub-RFC [RFC]: Bump vLLM-Omni Pin toded89346and Properly Fix the Async Rollout Server Lifecycle #492 §5.3's "wake must not self-resume"; the final [BREAKING][vllm_omni, rollout, trainer, ci, docker, diffusion, model, tests] fix: bump vllm-omni pin to 0.28 and fix the async rollout server lifecycle #497 commit made resume idempotent and kept it.resume_generation_replicas()— verl'sCheckpointEngineManager.update_weightsalready ends withresume_generation(verlcheckpoint_engine/base.py:506-558, resume at:556), andPPOTrainerColocateAsync.on_step_endcalls it again (:53) idempotently.omni_syncbridge (ray_omni_trainer.py:78-84, bothon_init_end+on_step_end→resume_generation_replicas) is retained as an init-time safety net for holds not preceded by a wake; its in-code TODO points at the rollout-side consolidation that [Bug] omni admission-resume is duplicated per trainer — own it on the rollout/weight-sync path instead #517 now proposes.2.3 Known limits at this pin (scope boundaries for PR 2 / PR 3)
:383-386; needs a vllm-omni fix + pin bump). Single-stage / thinker-only rollouts — which is exactly what [trainer, tests, recipe] feat: add omni colocate-async V1 trainer #337/[trainer, cfg, tests, recipe] feat: add omni separate-async V1 trainer #380 and the registered smoke do — are correct. PR 2 / PR 3 land thinker-only first; talker-in-the-loop is out of scope (§9).num_computed_tokens == 0, answers correctly, withenable_prefix_caching=True) passed. Keep that signature as a standing gate for PR 2/PR 3 GPU runs.AsyncOmni.reset_mm_cache()never finds the cache (upstream #6972, 2026-09-03); [BREAKING][vllm_omni, rollout, trainer, ci, docker, diffusion, model, tests] fix: bump vllm-omni pin to 0.28 and fix the async rollout server lifecycle #497 works around it renderer-side (:225-231, drop after #7003 lands and the pin moves).2.4 What transfers from verl — and what must not
AsyncLLM)AsyncOmni, post-#497)pause_generation(wait_for_inflight_requests=False, clear_cache=…)— pause is abort + token delivery + stay-paused (vllm_async_server.py:922-925)abort(ids)while generate is live, thenpause_generation(mode="abort", …); samereset_prefix_cachekwarg,{aborted_count, request_ids}return_resolve_sleep_level()may return 2 (:1116-1133)sleep()defaults to 2 andwake_upafter level-2 raisesNotImplementedError["kv_cache", "weights"](:1112-1114)["weights"](:198-199) — #337's["kv_cache","weights"]divergence was rejected; revisit only with GPU AR-KV evidencewake_upexplicitly resumes (:217); abort-path hold lifted by trainerresume_generation_replicaswait_for_requests_to_drain(300 s, raises)abort_all_requestsThe sync omni path shares the same server, so this contract covers
omni_syncand both async modes alike.3. Design
3.1 Core principle: minimal new code (unchanged)
OmniPPOTrainerSyncis 20 lines including the bridge hooks (ray_omni_trainer.py:65-84); its only real override is_init_tokenizer()— wiring tokenizer/processor fromOmniModelConfiginstead of the HF path. The async trainers are equally thin subclasses of verl'sPPOTrainerColocateAsync/PPOTrainerSeparateAsync; agent-loop lifecycle, weight sync, abort/resume, reward, advantage, loss are all inherited.Registration uses verl's
TRAINER_REGISTRY(trainer_base.py:1897); unregistered modes fail withValueError: Unknown trainer '<name>'(:1918-1924); registration is import-time, triggered bymain_omni.pyimportingverl_omni.trainer.omni(main_omni.py:39). Fresh names (omni_colocate_async,omni_separate_async) — duplicate names raise at import (:1906-1911).3.2 PR 2 —
OmniPPOTrainerColocateAsyncSubclass of
PPOTrainerColocateAsync(verltrainer_colocate_async.py:26); the only override is_init_tokenizer(identical toOmniPPOTrainerSync). Registered as"omni_colocate_async". #337 also carried anon_train_beginoverride readingv1.omni_colocate_async.num_warmup_batches— dropped in the rebase: warmup inherits verl's hook, which reads the genericv1.colocate_asynckey (§3.6).Inherited lifecycle (verified at pin
fefb080):abort_replicas→abort_all_requests(abort-then-pause, partial tokens preserved);sleep_replicas→ level-1 sleep;update_weightsends withresume_generation(verlbase.py:556) so the aborted hold is lifted before the next generate. UsesFullyAsyncLLMServerClientfor abort/retry — the agent loop is unaware of interruptions; aborted generations are re-submitted with accumulatedtoken_ids(§4, row 3). No new bridges, no server changes.3.3 PR 3 —
OmniPPOTrainerSeparateAsyncSubclass of
PPOTrainerSeparateAsync(verltrainer_separate_async.py:43). Overrides_init_tokenizer(identical) plus theparameter_sync_stepfix (§3.4). Registered as"omni_separate_async".Lifecycle — corrected to the pin. verl #7373 moved the hybrid↔standalone switching out of the sample hooks; the old RFC's
on_sample_begin → switch_to_rollout/on_sample_end → switch_to_trainermapping is stale (on_sample_begin/on_sample_endat:234-242are accounting-only). Actual hook wiring atfefb080:where
switch_to_rollout(:370-376) =update_weights + resume_generation_replicas + add_replicas_to_balancer+ mode flag, andswitch_to_trainer(:378-383) =remove_replicas_from_balancer + abort_replicas + sleep_replicas. Hybrid replicas join the standalone balancer viaadd_replicas_to_balancer(:385-390; sticky cache cleared on re-registration so load redistributes). The wholeHybridRolloutSwitchConfig/enable_switch(GPU lending) machinery is available and on by config. Note the standing guards: naive checkpoint-engine backend rejected (:63-65), reward model must use a resource pool (:66-71), rollout disaggregation incompatible withenable_switch(:79-86).Standalone vLLM-Omni pool reuses the auto-generated deploy config (from
OmniRolloutPipelineBase) pointed at standalone endpoints — no adapter changes; #380 carries the working example + docs.3.4
parameter_sync_stepfix (confirmed still needed)PPOTrainer.__init__readsself.parameter_sync_stepfromconfig.trainer.v1.get(trainer_mode, {}).get("parameter_sync_step", 1)(verltrainer_base.py:136) — fortrainer_mode="omni_separate_async"the key misses and silently yields 1, disabling Decoupled PPO (π_old stabilization via CPU save/restore,trainer_separate_async.py:157-178;train_batch_size == parameter_sync_step * ppo_mini_batch_sizeis asserted at:54-58against the hardcodedv1.separate_asynckey). Verified unfixed on verl main as of 2026-09-03 (the read sits attrainer_base.py:138there — two lines were added above it; the bug itself is untouched) — the override stays until an upstream fix lands:Which key wins:
v1.separate_async(the validated key) — theomni_separate_asyncconfig stub therefore omitsparameter_sync_stepto avoid a silently-ignored duplicate.3.5
OmniDetachActorWorkerverl's
DetachActorWorker(verl/experimental/separation/engine_workers.py:36; 169 lines, does not overrideinit_model) inherits verl'sActorRolloutRefWorker, whoseinit_modeldispatches via_target_/hydra-instantiate (omega_conf_to_dataclasswithoutdataclass_type,verl/workers/engine_workers.py:538-539; theHFModelConfigannotation is a stale hint). Both paths produce a realOmniModelConfigtoday, but that is coincidence, not guarantee — the detach path must route through verl-omni'sinit_modelso omni-specific behavior can never silently diverge.The shipped precedent is composition (MRO), not
__init__re-routing —DiffusionDetachActorWorker(ActorRolloutRefWorker, DetachActorWorker)atverl_omni/workers/detach_actor_worker.py:43, with CPU-snapshot save/restore at:66-84. PR 3 adds the omni twin the same way (matching #380's actual hunks inworkers/engine_workers.py/workers/omni_engine_workers.py):verl's save/restore only touches
self.actor.engine.moduleandself.config.actor.strategy(handler dispatch coversfsdp/fsdp2/ megatron,engine_workers.py:73-108) — both set identically by verl-omni's__init__, so all ofDetachActorWorker's Decoupled-PPO logic is reused verbatim. Strategy dispatch covers"fsdp"/"fsdp2"→OmniFSDPEngine.OmniPPOTrainerSyncalready runs through verl-omni'sinit_model, so this is consistent with current omni behavior — no regression.3.6 Config: no new stubs — the generic keys are the single read path
The generic stubs already exist via generation from verl's base (
_generated_omni_trainer.yaml:1021-1045:colocate_async: {num_warmup_batches: 1},separate_async: {num_warmup_batches: 1, parameter_sync_step: 4, hybrid_rollout: {...}}), and every consumer reads them: colocate warmup (verltrainer_colocate_async.py:43), separate-async warmup and validation (trainer_separate_async.py:53-58,:199), and — via §3.4 —parameter_sync_step.trainer_modeis a registry string, not a config key: nothing readsv1.omni_*.PR 2 / PR 3 therefore add no
omni_*config stubs. Anomni_colocate_async: {num_warmup_batches: ...}stub would be silently-ignored dead config — the same misleading-duplicate hazard §3.4 avoids forparameter_sync_step. One rule covers both modes: generic keys are the single read path (this is why #337's omni-key warmup override and its stub are dropped in the rebase, §3.2). Validtrainer_modevalues:omni_sync,omni_colocate_async,omni_separate_async; unknown names fail fast inget_trainer_cls(:1918-1924).3.7 Bug fix:
adapter_nameforwarding (still open)OmniFSDPEngine.get_per_tensor_param(omni_impl.py:58) accepts**kwargsbut does not forwardadapter_nametocollect_lora_params(call site:77-81), which does acceptadapter_name: str = "default"(fsdp_utils.py:422-429). Non-default LoRA adapters would broadcast wrong weights in separate_async. Fix (~3 lines) + CPU routing test ship in PR 3. (UpstreamFSDPEngineWithLMHeadhas the same latent bug — out of scope here.)4. Verified compatibility (post-#497 evidence)
token_ids == [7,8,9]survives),{"aborted_count", "request_ids"}shape —test_vllm_omni_async_server_lifecycle_on_cpu.py:115/:150; empty-list abort still pauses:133"aborted"(vllm_omni_strategy_base.py:225-232); verl client resumes fromprompt_ids + token_idsonstop_reason in ("aborted","abort")(verlllm_server.py:404-455) — frozen seam, do not touchlevel=1+ keywordtags=+ ACK fail-closed + LoRA-cache invalidation, uniform across HYBRID/COLOCATED — lifecycle tests:171/:189/:299/:323:217, tests:536/:546/:557); AR-only engines need no trainer resume after sleep/wakeupdate_weightsordering endsresume_generation(verlbase.py:506-558, resume at:556); colocateon_step_endresume is idempotent (:53)vLLMOmniHttpServerasray.remoteactor;GlobalRequestLoadBalancer.add_servers/remove_servers+ sticky-cache invalidation (verlllm_server.py:92-95/:123-149/:159-178)AgentLoopWorker._compute_multi_modal_inputs(verlagent_loop.py:838-894) → TransferQueue postprocess (agent_loop_tq.py:150-228)input_idstensors — no re-tokenization (unchanged since #232 series)OmniDetachActorWorker+OmniFSDPEngineDiffusionDetachActorWorker(detach_actor_worker.py:43); DetachActorWorker save/restore surface verified at pin (engine_workers.py:123-161)OmniCheckpointEngineManager(checkpoint_engine.py:19) inherits; both sides land onadapter_name="default"— non-default is what §3.7 fixestest_omni_sync_resume_bridge_on_cpu.py:75/:82/:89/:109/:128); AVQA GSPO 240-step GPU run green, +10-12% throughput (#497)5. Risks & mitigations
trainer_separate_async.py:315/:324)update_weightstail already cover them). #380's worker hunks are the overlap zone — settle #517's placement before PR 3's final review; PR 2 is unaffectedmax_concurrency(control RPCs queued behind generates)_CONTROL_CONCURRENCY_GROUPthere, not here)enable_prefix_caching=TrueGPU signature as a standing gate; recipes set prefix caching explicitly (AR default silently True)parameter_sync_stepsilently 1trainer_base.py:136upstreamadapter_nameforwardingverl/workers/rollout/router.py), #7227 (delta-sharded sync)GlobalRequestLoadBalancerfromllm_serveradapts then6. Implementation plan & coordination
6.1 PR 2 —
omni_colocate_async(rebase #337; owner @zhtmike)#337 (draft, last push 2026-08-14, pre-bump) splits cleanly:
ray_omni_trainer_colocate_async.py(registration +_init_tokenizeronly),__init__.pyexport, MMK12 example script, CPU tests (registration / tokenizer wiring).engine.abort_with_output_ids(...)(nonexistent atded89346), sets wake tags["kv_cache","weights"](rejected — see §2.4), retains an_enqueue_abort_outputfallback on the success path, and adds_CONTROL_CONCURRENCY_GROUP(salvage into [RFC][tracking] Refactorvllm_omni_async_server.py— Align with Upstream verl, Remove Temp Patches #377).on_train_beginoverride readingv1.omni_colocate_async, itsomni_trainer.yamlstub, and the omni-key warmup test — superseded by the generic-key rule (§3.6).Gate (per #491 §6): MMK12-style colocate run — abort+sleep boundary observed, step time without drain compensation (the drain knob no longer exists; 120 s/step stalls would indicate a real bug),
mm_processor_cache_gb=0.6.2 PR 3 —
omni_separate_async(rebase #380; owner @AndyZhou952)#380 (draft, updated 2026-09-01, +907/−15 over 14 files) needs no server changes — its three historical engine blockers are fixed. Rebase keeps: trainer (+ export),
workers/engine_workers.py/workers/omni_engine_workers.py/engine/fsdp/omni_impl.py(detach wiring + §3.7), docs (separate_async_omni.md), example + smoke scripts, 4 CPU tests. Two open threads from its discussion: convergence parity and "actor-updating takes extensive time" — the latter was the old frontend-only-pause symptom, expected to disappear against the new pin (verify in the gate).Gate (per #491 §6 / #399): hard-abort/weight-sync gate (partial-token continuation, no stale/late outputs) + a convergence run with correct
parameter_sync_step(Decoupled PPO engaged, not silently 1).6.3 #517 coordination (open design question)
#517 (2026-09-03, @knlnguyen1802) proposes moving resume ownership onto the rollout/weight-sync path and deleting the per-trainer bridges (including the duplicated pattern in diffusion
trainer_separate_async.py:315/:324and, eventually, theomni_syncsafety net). Neither PR 2 nor PR 3 should grow bridge code that #517 would immediately delete: colocate needs none (wake self-resume +update_weightstail suffice), separate inherits verl'sswitch_to_rolloutresume. Sequence: land PR 2/PR 3 bridge-free; #517 consolidates server-side; then delete theomni_syncsafety net.6.4 Sequencing (mirrors #491 §4)
Both PRs depend only on landed work (1A/1B ✅, 1C ✅) and can proceed in parallel after rebase. Post-stack follow-ups (not here): #452 (unblocked — upstream #6645 merged 2026-09-01; open non-draft after), #377 remainder (concurrency group), #391, #388 B4, #6442 close-out.
Files changed:
verl_omni/trainer/omni/ray_omni_trainer_colocate_async.pyverl_omni/trainer/omni/ray_omni_trainer_separate_async.pyverl_omni/trainer/omni/__init__.pyomni_*stubs; the generic keys are the read path (§3.6)verl_omni/workers/engine_workers.py/omni_engine_workers.pyOmniDetachActorWorker(§3.5)verl_omni/workers/engine/fsdp/omni_impl.pyadapter_nameforwarding (~3)examples/gspo_trainer/qwen3_omni/run_*_colocate_async.shexamples/separate-async standalone+hybrid scripts,docs/separate_async_omni.mdtests/trainer/omni/test_omni_*_async_on_cpu.py7. Verification
CPU (pattern:
tests/**/*_on_cpu.py,--asyncio-mode=auto):get_trainer_cls("omni_colocate_async"/"omni_separate_async")resolve; unknown mode raisesValueError._init_tokenizerwiring (processor + tokenizer fromOmniModelConfig) for both trainers.v1.colocate_async.num_warmup_batches(no override; anomni_colocate_asynckey, if present, is ignored).parameter_sync_stepread from the validatedv1.separate_asynckey (mock config where the keys differ; assert not silently 1).adapter_namerouting:get_per_tensor_param(adapter_name="old")reachescollect_lora_params.OmniDetachActorWorker: routes through omniinit_model(isinstance(worker.actor.engine, OmniFSDPEngine),model_configisOmniModelConfig);save_model_to_cpu/restore_model_from_cpuround-trip.trainer_modevalues resolve end-to-end with only the generic stubs present (noomni_*keys required).GPU gates (per #491 §6): PR 2 — MMK12-style colocate run (§6.1). PR 3 — hard-abort/weight-sync gate + convergence run (§6.2). Standing regressions: #6442 prefix signature with explicit
enable_prefix_caching=True; CUDA-free server PID after wake. Scope: thinker-only/single-stage AR (§2.3.1).8. Future work
omni_syncinit-time bridge.wait_for_requests_to_drainfor real when DP is supported (replacing the documented no-op,:354-356)._reset_frontend_mm_cache(:225-231).parameter_sync_stepfix: verltrainer_base.py:136— once fixed, delete the §3.4 override.vllm_omni_async_server.py— Align with Upstream verl, Remove Temp Patches #377 remainder: dedicated control-RPC concurrency group.9. Out of scope
"aborted"stop-reason seam or the strategy layer's request envelope ([RFC] Split vLLM-Omni generation strategies and define typed multimodal rollout I/O #403 series owns it).10. Tracking checklist
omni_colocate_asynctrainer (rebase [trainer, tests, recipe] feat: add omni colocate-async V1 trainer #337; server hunk + omni-key warmup override dropped — superseded by [BREAKING][vllm_omni, rollout, trainer, ci, docker, diffusion, model, tests] fix: bump vllm-omni pin to 0.28 and fix the async rollout server lifecycle #497 / §3.6)mm_processor_cache_gb=0omni_separate_asynctrainer (rebase [trainer, cfg, tests, recipe] feat: add omni separate-async V1 trainer #380; §3.4 param-sync fix, §3.5OmniDetachActorWorker, §3.7 adapter fix)enable_prefix_caching=True), CUDA-free PID after wakevllm_omni_async_server.py— Align with Upstream verl, Remove Temp Patches #377 remainder, [RFC][tracking] Fail Fast — No Silent Fallbacks #388 B4, [RFC] Stable vLLM-Omni Rollout Interface for Diffusion RL #391 sequencingRelated: #491 (umbrella — PR 2/PR 3 rows), #497 (landed engine foundation), #492 (sub-RFC, closed), #337 / #380 (landing vehicles), #517 (resume ownership), #399 / #413 (finalize against the bumped pin), #6442 / #6972 / #7003 (upstream), #445 (patch ledger).