Skip to content

flash mla kernal是否支持同一batch下不同query的动态token数? #190

Description

@echo-timeless

q: (batch_size, seq_len_q, num_heads_q, head_dim)
flash mla算子的输入这个seq_len_q是定死的,那么是否不支持不同的seq_len_q进入kernal计算?

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions