ArcticInference is fantastic work for accelerating LLM inference. We integrated the suffix feature into the vLLM-Ascend framework on NPU platforms with excellent performance outcomes.
We noticed ArcticInference includes the LSTM speculative decoding method, which has drawn lots of attention from the vLLM community. So together with experts from the vLLM-Ascend community, we’ve adapted the ArcticInference repository to run on NPU platforms, with full support for the LSTM speculative decoding method. I would like to contribute the fully validated adaptation changes back to the upstream community in the form of the ArcticInference-Ascend repository. Of course, we will maintain this repository going forward.
I'd like to ask for the your team's advice on whether we can host the ArcticInference-Ascend repository within the snowflakedb organization as snowflakedb/ArcticInference-Ascend, for everyone in the community to use. @sfc-gh-aqiao @sfc-gh-yewang please cc.
ArcticInference is fantastic work for accelerating LLM inference. We integrated the suffix feature into the vLLM-Ascend framework on NPU platforms with excellent performance outcomes.
We noticed ArcticInference includes the LSTM speculative decoding method, which has drawn lots of attention from the vLLM community. So together with experts from the vLLM-Ascend community, we’ve adapted the ArcticInference repository to run on NPU platforms, with full support for the LSTM speculative decoding method. I would like to contribute the fully validated adaptation changes back to the upstream community in the form of the ArcticInference-Ascend repository. Of course, we will maintain this repository going forward.
I'd like to ask for the your team's advice on whether we can host the ArcticInference-Ascend repository within the snowflakedb organization as snowflakedb/ArcticInference-Ascend, for everyone in the community to use. @sfc-gh-aqiao @sfc-gh-yewang please cc.