Skip to content

Qwen3-VL在bm1688 SOC 上文字生成速度缓慢,仅有0.5 tokens/s #137

Description

@zhangjingxian1998

标准示例:

环境:

soc环境
transformers:4.57.1
torch:2.9.0+cpu
LLM-TPU:67a836d 2025.10.30
tpu-mlir:504f27a 2025.10.26
driver版本:0.4.9
sdk版本0.4.9
libsophon:#40 SMP Sun Nov 24 22:38:19 CST 2024

路径:

/home/linaro/LLM-TPU-main/models/Qwen3_VL/python_demo

操作:

python3 pipeline.py -m /data/qwen3-vl-4b-instruct_w4bf16_seq2048_bm1688_2core_20251026_141708.bmodel -c ../config

问题:

在bm1688 soc平台上推理Qwen3-VL 生成token缓慢,完全达不到示例上给的13 token/s左右。使用的模型是提供转换好的bmodel。
推理过程中,tpu利用率不高,仅有15%左右。是否有办法设置tpu利用率?
Image

Image Image

其他:

如果是自己编译的模型,需要注明使用拉下来的模型,能否跑通

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions