Currently, async scheduling is only supported with EAGLE/MTP kind of speculative decoding https://github.com/vllm-project/vllm/blob/c4e744dbd41f23ee7fd554a86cc1bf516082552d/vllm/config/vllm.py#L574-L576
Currently, async scheduling is only supported with EAGLE/MTP kind of speculative decoding
https://github.com/vllm-project/vllm/blob/c4e744dbd41f23ee7fd554a86cc1bf516082552d/vllm/config/vllm.py#L574-L576